Text search method and device based on memory index and memory mapping

CN122711751APending Publication Date: 2026-09-08NANCHANG XINGWEI SOFTWARE DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610963831.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

然而,现有技术在单次查询时需多次执行文件定位与读取操作,产生大量随机磁盘I/O(Input/Output,输入输出)开销,查询响应延迟较高,在批量查询场景下性能下降尤为明显

Benefits of technology

[0008]This application provides a text search method and apparatus based on memory indexing and memory mapping. First, the storage ranges of the index area and data area are clearly defined based on the attribute information of the text library file, realizing the partition boundary between index data and content data within the text library, and providing accurate addressing basis for the differentiated processing of the two types of data. Then, all index information of the index area is read into memory at once to form an index dataset, so that all subsequent search operations for target identifiers can be completed directly in memory, without repeatedly accessing the index area on the disk during each query, avoiding the cumulative I/O overhead caused by multiple disk file locations and fragmented reads. At the same time, the entire data area is mapped to the virtual address space of the current process, so that the text content stored on the disk can be read in the form of memory address access, without initiating disk read operations through the file system interface. Finally, when a text query request is received, the location description information corresponding to the target identifier is first matched in the index dataset on the memory side, and then the corresponding text content is read directly from the virtual address space and returned in combination with the mapped address, so that the entire query execution process is completed through memory access. Therefore, this application loads all index information of the index area into memory to form an index dataset and maps the data area to the process virtual address space, so that text identification query and content reading are both completed based on memory operations, fundamentally eliminating the performance overhead caused by frequent random disk I/O, reducing the response latency of text search, and thus improving the running efficiency of text content search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122711751A_ABST
    Figure CN122711751A_ABST
Patent Text Reader

Abstract

The application discloses a text searching method and device based on memory index and memory mapping, comprising: obtaining a text library file to be searched; determining a first storage range of an index area in the text library file and a second storage range of a data area in the text library file according to attribute information of the text library file; reading all index information contained in the index area into the memory according to the first storage range, and forming an index data set for searching in the memory; mapping the data area to a virtual address space of a current process according to the second storage range, and obtaining a mapping address of the data area in the virtual address space; in response to a text query request containing a target identifier, searching for position description information corresponding to the target identifier in the index data set; and reading and returning target text content from the virtual address space according to the position description information and the mapping address. The application can reduce the response delay of text searching and improve the text content searching efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, specifically to a text search method and apparatus based on memory indexing and memory mapping. Background Technology

[0002] With the rapid development of automotive electronics technology, automotive diagnostic software needs to quickly parse the corresponding readable text content based on the text identifiers returned by the diagnostic protocol, thereby supporting core diagnostic functions such as fault code interpretation and data stream display. The response speed of text search directly affects the operating efficiency and user experience of the diagnostic software.

[0003] Currently, the industry commonly uses disk-based text libraries for text retrieval. These libraries consist of an index area sorted by text identifiers and a data area storing the text content sequentially. During a query, a search algorithm first locates the position information corresponding to the target identifier in the index area, and then reads the corresponding text content from the data area based on that position information. However, existing technologies require multiple file location and read operations per query, generating significant random disk I / O (Input / Output) overhead, resulting in high query response latency. Performance degradation is particularly pronounced in batch query scenarios. In summary, existing technologies suffer from high disk I / O overhead and low query efficiency, making it difficult to meet the real-time operational requirements of diagnostic software.

[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention

[0005] This application provides a text search method and apparatus based on memory indexing and memory mapping, which can reduce the response latency of text search and improve the efficiency of text content search.

[0006] In a first aspect, embodiments of this application provide a text search method based on memory indexing and memory mapping, including: Obtain the text library file to be searched, the text library file including an index area and a data area; Based on the attribute information of the text library file, determine the first storage range of the index area in the text library file, and determine the second storage range of the data area in the text library file; According to the first storage range, all index information contained in the index area is read into memory to form an index dataset for searching in memory; Based on the second storage range, the data area is mapped to the virtual address space of the current process to obtain the mapped address of the data area in the virtual address space; In response to a text query request containing a target identifier, the location description information corresponding to the target identifier is searched in the index dataset; Based on the location description information and the mapping address, the target text content is read from and returned from the virtual address space.

[0007] Secondly, embodiments of this application provide a text search device based on memory indexing and memory mapping, comprising: The file acquisition module is used to acquire the text library file to be searched, the text library file including an index area and a data area; The information determination module is used to determine, based on the attribute information of the text library file, a first storage range of the index area in the text library file and a second storage range of the data area in the text library file; The index loading module is used to read all the index information contained in the index area into memory according to the first storage range, so as to form an index dataset for searching in memory; The mapping establishment module is used to map the data area to the virtual address space of the current process according to the second storage range, and obtain the mapping address of the data area in the virtual address space; The query response module is used to respond to a text query request containing a target identifier and search for location description information corresponding to the target identifier in the index dataset; The content reading module is used to read and return target text content from the virtual address space based on the location description information and the mapping address.

[0008] This application provides a text search method and apparatus based on memory indexing and memory mapping. First, the storage ranges of the index area and data area are clearly defined based on the attribute information of the text library file, realizing the partition boundary between index data and content data within the text library, and providing accurate addressing basis for the differentiated processing of the two types of data. Then, all index information of the index area is read into memory at once to form an index dataset, so that all subsequent search operations for target identifiers can be completed directly in memory, without repeatedly accessing the index area on the disk during each query, avoiding the cumulative I / O overhead caused by multiple disk file locations and fragmented reads. At the same time, the entire data area is mapped to the virtual address space of the current process, so that the text content stored on the disk can be read in the form of memory address access, without initiating disk read operations through the file system interface. Finally, when a text query request is received, the location description information corresponding to the target identifier is first matched in the index dataset on the memory side, and then the corresponding text content is read directly from the virtual address space and returned in combination with the mapped address, so that the entire query execution process is completed through memory access. Therefore, this application loads all index information of the index area into memory to form an index dataset and maps the data area to the process virtual address space, so that text identification query and content reading are both completed based on memory operations, fundamentally eliminating the performance overhead caused by frequent random disk I / O, reducing the response latency of text search, and thus improving the running efficiency of text content search. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is an application environment diagram of the text search method based on memory indexing and memory mapping provided in the embodiments of this application; Figure 2 This is a flowchart illustrating the text search method based on memory indexing and memory mapping provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of the text search device based on memory indexing and memory mapping provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0011] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of systems and methods consistent with those detailed in the appended claims or with some aspects of this application.

[0012] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover descriptions such as non-exclusive inclusion, so that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.

[0013] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0014] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.

[0015] To address the aforementioned technical problems, this application provides a text search method and apparatus based on memory indexing and memory mapping, which can reduce the response latency of text search and improve the efficiency of text content search.

[0016] Figure 1 This is a diagram illustrating the application environment of a text search method based on memory indexing and memory mapping in one embodiment. (Refer to...) Figure 1This text search method based on memory indexing and memory mapping is applied to a text search system based on memory indexing and memory mapping. The system includes a terminal 110 and a server 120. The terminal 110 and server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal; a mobile terminal can be at least one of a mobile phone, tablet, or laptop. The server 120 can be a standalone server or a server cluster consisting of multiple servers. Server 120 is configured to execute the aforementioned text search method based on memory indexing and memory mapping, including: obtaining a text library file to be searched, the text library file including an index area and a data area; determining a first storage range of the index area in the text library file and a second storage range of the data area in the text library file based on the attribute information of the text library file; reading all index information contained in the index area into memory according to the first storage range to form an index dataset for searching in memory; mapping the data area to the virtual address space of the current process according to the second storage range to obtain the mapping address of the data area in the virtual address space; responding to a text query request containing a target identifier, searching for location description information corresponding to the target identifier in the index dataset; and reading and returning the target text content from the virtual address space according to the location description information and the mapping address.

[0017] Please see Figure 2 , Figure 2 This is a flowchart illustrating a text search method based on memory indexing and memory mapping according to an embodiment of this application. This embodiment primarily uses the application of this text search method based on memory indexing and memory mapping to a server as an example for illustration. Specifically, the text search method based on memory indexing and memory mapping provided in this embodiment may include the following steps: S1. Obtain the text library file to be searched. The text library file includes an index area and a data area. Specifically, for step S1, the target text library file to be searched is first obtained. The target text library file is a pre-built structured storage file, which is internally divided into two contiguous and functionally independent storage areas at the physical storage level: an index area and a data area. The index area is used to carry the position association data between all text identifiers and their corresponding text content, while the data area is used to carry the raw byte data of all text content. Taking the automotive diagnostic scenario as an example, when the diagnostic software starts, it will obtain the diagnostic text library file for the corresponding vehicle model or protocol from the locally specified resource storage path. This file has been organized according to the rule of separating the index and content, and can be directly used for subsequent text parsing and querying.

[0018] S2. Based on the attribute information of the text library file, determine the first storage range of the index area in the text library file, and determine the second storage range of the data area in the text library file; Specifically, in step S2, the attribute information inherent in the text library file is read and parsed. Based on the region location parameters recorded in the attribute information, the first storage range corresponding to the index area in the entire text library file and the second storage range corresponding to the data area in the text library file are determined. The storage range is defined by the starting storage offset of the region and the total length of the region, which can completely delineate the physical storage boundary of the corresponding functional area in the file. The storage ranges of the two regions do not overlap and remain continuous. This step, by completing the physical boundary division of the index data and content data, provides a precise addressing basis for subsequent differentiated loading and mapping operations for the two regions, avoiding problems such as out-of-bounds reading or reading data from irrelevant regions. By clearly defining the accurate storage range of the two regions, subsequent operations can directly locate the start and end positions of the target region without having to additionally traverse and identify region boundaries in the file, ensuring the accuracy of subsequent operations.

[0019] S3. Based on the first storage range, read all the index information contained in the index area into memory to form an index dataset for searching in memory; Specifically, for step S3, based on the determined first storage range, the starting position of the index area in the text library file is located, and all the index information contained in the index area is read from the disk file into the pre-allocated system memory space in a continuous batch reading manner, forming an index dataset in memory that can be directly accessed and used to support search operations; this index dataset completely retains the association relationship between all text identifiers and corresponding position description information in the original index area, and is consistent with the index area data on the disk.

[0020] This step migrates the access medium for index data from disk to memory, fundamentally avoiding repeated accesses to the disk index area during the query process. This loading operation is completed in one go during the initialization phase, using a continuous batch read approach to access the disk. Compared to fragmented random reads during the query process, single continuous reads offer higher disk I / O efficiency. Furthermore, after loading, all subsequent searches for text identifiers are performed directly in memory. Regardless of the number of subsequent query requests, the disk read operation for the index area is executed only once, reducing the overall frequency of disk access.

[0021] S4. Based on the second storage range, map the data area to the virtual address space of the current process to obtain the mapping address of the data area in the virtual address space; Specifically, in step S4, based on the determined second storage range, the memory mapping mechanism provided by the operating system kernel is invoked to map the entire data area of ​​the text library into the virtual address space of the currently running process, and the starting mapping address corresponding to this mapped area is obtained. This mapping operation establishes a one-to-one correspondence between the file data in the data area on the disk and the virtual memory address of the process. After the mapping is completed, the process's read operation on this virtual address segment will be automatically converted by the operating system kernel to access the corresponding file content. This changes the text content reading mode from disk file I / O reading to memory address access, eliminating the need to actively initiate disk read operations during the query process. The mapping establishment process only completes the binding of the address mapping relationship and does not load all data into physical memory at once. While saving memory resources, it achieves the effect of reading disk file content in the form of memory access, balancing resource consumption and read performance.

[0022] S5. In response to a text query request containing a target identifier, search for location description information corresponding to the target identifier in the index dataset; Specifically, in step S5, when a text query request carrying a target identifier is received, a matching search is performed within the in-memory index dataset using the target identifier as the search keyword to locate the location description information corresponding to the target identifier. This location description information indicates the relative storage location and corresponding data range of the target text content in the data area, serving as the basis for subsequent text content retrieval. The rapid matching of the target based on the in-memory index dataset provides precise location guidance for subsequent text content retrieval. For example, when diagnostic software needs to parse the content corresponding to a certain diagnostic text identifier, it will pass that identifier as the target identifier to the query interface, match the location description information corresponding to that identifier in the in-memory index dataset, and determine the location and data size of the corresponding text content.

[0023] S6. Based on the location description information and the mapping address, read and return the target text content from the virtual address space; Specifically, for step S6, combining the previously obtained location description information and the mapping address of the data area, the absolute starting address of the target text content in the virtual address space is calculated by first using the mapping address as the base address and superimposing the relative position parameters in the location description information. Then, starting from this starting address, the original data of the target text content of the corresponding length is read, and the read text content is organized and returned to the request initiator. The text content is read directly from the virtual address space according to the memory mapping mechanism, without needing to call the file system interface to initiate disk I / O operations. The entire reading process is essentially a data copying operation of the memory address, with a reading speed on the same order of magnitude as conventional memory access, higher than the speed of disk file reading, achieving extremely low query response latency.

[0024] This embodiment loads all index information from the index area into memory to form an index dataset and maps the data area to the process virtual address space during the initialization preprocessing stage. This allows the two core query steps of identifier lookup and content reading to be completed at the memory level. This eliminates the need to repeatedly perform disk file location and data reading operations during each query, fundamentally eliminating the performance overhead caused by frequent random disk I / O, reducing the response latency of text search, stably supporting high-frequency text query needs, and improving the running efficiency of text content search and the smoothness of system response.

[0025] Furthermore, in some embodiments, step S2, "determining the first storage range of the index area in the text library file and determining the second storage range of the data area in the text library file based on the attribute information of the text library file," may specifically include: S21. Read the file header information of the text library file, and extract the starting position of the index area in the text library file and the total number of index entries contained in the index area from the file header information. The attribute information includes the file header information. Specifically, for step S21, the file header information of the text library file is first read. This file header information belongs to the attribute information of the text library file and is a structured metadata area pre-written at the beginning of the file, used to record the overall structure configuration parameters of the file. The corresponding fields in the file header are parsed to extract two core parameters: one is the starting position of the index area in the entire text library file, that is, the byte offset of the first byte of the index area relative to the beginning of the file; the other is the total number of index entries contained in the index area, that is, the total number of index entries corresponding to the text identifiers stored in the index area.

[0026] For example, in a text library file for automotive diagnostics, the file header area occupies a fixed 64 bytes of storage space at the beginning of the file. Bytes 8 to 11 record the starting position of the index area, and bytes 12 to 15 record the total number of index entries. After reading the file header, it can be parsed that the starting position of the index area is 64 bytes off the file, and the total number of index entries is 100,000.

[0027] S22. Based on the starting position and the total number of index entries, and combined with the preset storage length occupied by a single index entry, calculate the ending position of the index area, and limit the first storage range with the starting position and the ending position; Specifically, in step S22, based on the extracted starting position of the index area and the total number of index entries, combined with the pre-agreed storage length occupied by a single index entry, the ending position of the index area is calculated. The storage length of a single index entry is a preset fixed number of bytes, and all index entries use a uniform storage format and length. Therefore, the total storage length of the index area can be calculated as "total number of index entries × storage length of a single index entry". Then, using the starting position of the index area as a reference, and adding the total storage length of the index area, the file offset corresponding to the last byte of the index area can be obtained, i.e., the ending position. Finally, the starting position and the ending position together define the first storage range corresponding to the index area.

[0028] For example, if a single index entry occupies a fixed 12 bytes of storage space, and the total number of index entries is 100,000, then the total storage length of the index area is 100,000 × 12 = 1,200,000 bytes. Combining the starting position of the index area of ​​64 bytes, the ending position of the index area is calculated to be 64 + 1,200,000 = 12,000,64 bytes. That is, the first storage range is a continuous storage area from the offset of 64 bytes to 120,0064 bytes within the file.

[0029] S23. Obtain the starting position of the data area in the text library file and the total length of the data area, calculate the ending position of the data area based on the starting position and the total length of the data, and limit the second storage range with the starting position and the ending position; Specifically, for step S23, the starting position of the data area in the text library file and the total length of the data area are obtained. These two parameters can also be parsed from the file header information. Then, based on the starting position of the data area, the total length of the data is added to calculate the ending position of the data area. Finally, the starting position and the ending position together define the second storage range corresponding to the data area. The data area is a contiguous storage region in the file, and its starting position usually connects with the ending position of the index area, forming a file structure in which the index area and the data area are arranged continuously.

[0030] For example, if the starting position of the data area is found to be 1200064 bytes from the file header information, and the total length of the data area is 5,000,000 bytes, then the ending position of the data area is 1200064 + 5,000,000 = 6200064 bytes. That is, the second storage range is a continuous storage area from offset 1200064 bytes to 6200064 bytes within the file.

[0031] This embodiment derives and calculates the complete storage range of the index area and data area by parsing the metadata parameters in the file header and combining them with the preset fixed length rules of the index items. It can accurately and efficiently define the physical boundaries of the two functional areas. The parameter source is reliable and the calculation logic is simple and controllable. It provides an accurate addressing basis for subsequent index loading and data area mapping operations, ensuring the accuracy and execution efficiency of partitioning.

[0032] Furthermore, in some embodiments, step S3, "reading all index information contained in the index area into memory according to the first storage range, so as to form an index dataset for searching in memory," may specifically include: S31. Based on the starting position and storage length defined by the first storage range, position the file read pointer to the starting position of the index area; Specifically, in step S31, based on the starting position and storage length parameters defined by the first storage range, the file system's pointer positioning interface is invoked to precisely move the current file's read pointer to the starting position of the index area, which is the offset position of the first byte of the index area in the text library file. This positioning operation ensures that the starting point of subsequent read operations is completely aligned with the physical starting point of the index area, avoiding the reading of invalid data in non-indexed areas such as the file header, and guaranteeing the accuracy of the read data range.

[0033] For example, if the first storage range limits the starting position of the index area to the 64th byte in the file and the total storage length is 1,200,000 bytes, the file read pointer will be jumped from the default file start position to the 64th byte position, so that the read start point coincides with the first byte of the index area, thus preparing the position for subsequent continuous reading.

[0034] S32. Using the storage length as the total reading volume, continuously read all the index information contained in the index area from the text library file into the memory buffer; Specifically, in step S32, using the storage length corresponding to the first storage range as the total reading volume, starting from the already located starting position, all index information contained in the index area is completely read from the text library file on the disk into a pre-allocated memory buffer through a single continuous read operation. The capacity of this memory buffer matches the storage length of the index area and is used to temporarily hold the raw binary index data read from the disk. By using a single continuous batch read method to load index data, only one disk I / O operation is triggered to complete the loading of all index data. Compared with fragmented multiple random reads, this significantly reduces the total disk I / O overhead and results in higher data loading efficiency.

[0035] For example, for an index area of ​​1,200,000 bytes, starting from byte 64, 1,200,000 bytes of raw data are read continuously at once and written completely into a pre-allocated memory buffer of the corresponding size, thus completing the migration of index data from disk to memory.

[0036] S33. Parse each index information in the memory buffer, and use the text identifier in each index information as the search key to establish a mapping relationship between each text identifier and the corresponding offset information and content length information, and use the mapping relationship as the index dataset; Specifically, in step S33, the raw binary index data stored in the memory buffer is parsed one by one to identify the text identifier, offset information, and content length information contained in each index information. Then, using the text identifier in each index information as the search keyword, a one-to-one mapping relationship is established between each text identifier and its corresponding offset information and content length information. Finally, all mapping relationships are integrated to form an index dataset that can be directly used for querying, completing the transformation from raw binary data to structured search data, and transforming the disk data that originally only had storage attributes into a memory data structure that can support fast matching and searching.

[0037] For example, each 12 bytes in the buffer corresponds to a complete index information, where the first 4 bytes are text identifiers, the middle 4 bytes are offset information, and the last 4 bytes are content length information. After splitting and parsing each text identifier according to a fixed length, the corresponding offset and length parameters are bound to each text identifier. All the relationships together form the index dataset, which resides in memory for subsequent query calls.

[0038] This embodiment uses a three-step process—precise pointer positioning, single continuous batch reading, and structured parsing mapping—to efficiently and accurately migrate the index data on the disk to memory and transform it into a structured index dataset that can be directly used for searching. This completes the index loading with minimal disk I / O overhead, providing memory-side data support for subsequent fast text queries and ensuring the efficiency of the index loading process and the availability of the data.

[0039] Furthermore, in some embodiments, step S33, "establishing a mapping relationship between each text identifier and its corresponding offset information and content length information," may specifically include: S331A. Arrange the parsed index information in order of the size of the text identifier, and construct an ordered array structure in memory. Each array element in the ordered array structure includes the text identifier and the offset information and content length information corresponding to the text identifier. Specifically, in step S331A, all the index information obtained from parsing is arranged sequentially according to the numerical order of the text identifiers, forming an ordered array structure in memory. Each element of this ordered array is an independent structured data unit, fully storing the text identifier of the corresponding entry, as well as the offset information and content length information corresponding to that text identifier; all elements in the array are strictly arranged in ascending order according to the numerical order of the text identifiers, maintaining a continuous and ordered storage state overall.

[0040] Since the original index data can be pre-generated and written to a file according to sorting rules, the elements can be filled directly in the order of reading when building the array, or sorting verification can be performed after loading to ensure orderliness; the array uses contiguous memory blocks for storage, and the element position corresponds one-to-one with the array index, which has good memory access locality and can quickly access the corresponding element through the index.

[0041] S332A. Based on the ordered array structure, establish a binary search path based on text identifier comparison. The binary search path is used to locate the array element corresponding to the target identifier by successively dividing the search interval in half when the target identifier is received. Specifically, for step S332A, based on the completed ordered array structure, a binary search path based on the comparison of text identifier values ​​is established. This search path is a standardized search execution logic: when a query request carrying a target identifier is received, the range of possible target identifiers is continuously narrowed by successively dividing the current search interval in half, and finally the array element corresponding to the target identifier is accurately located.

[0042] The specific execution logic is as follows: When the search starts, the first and last indices of the entire array are used as the left and right boundaries of the initial search interval; each time, the middle index position of the current interval is taken, and the text identifier of the element at that position is compared with the target identifier: if the two values ​​are equal, the location is successful, and the offset information and content length information within the element can be directly extracted; if the target identifier value is smaller, the search interval is reduced to the left half of the middle position; if the target identifier value is larger, the search interval is reduced to the right half of the middle position; the above process of splitting and comparing is repeated until the target element is found, or it is confirmed that the search interval is invalid or the target identifier does not exist.

[0043] This embodiment constructs an ordered array structure sorted by text identifiers and establishes a binary search path to form an efficient and reliable index search system in memory. It has the advantages of simple data structure and compact memory layout, and can maintain stable logarithmic search efficiency in large-scale index data scenarios.

[0044] Furthermore, in some embodiments, step S33, "establishing a mapping relationship between each text identifier and its corresponding offset information and content length information," may specifically include: S331B. Using the text identifiers in each parsed index information as hash keys, and the corresponding offset information and content length information as hash values, a hash mapping table structure is constructed in memory. Specifically, in step S331B, all parsed index information is processed one by one. Using the text identifier in each index information as the hash key, the offset information and content length information corresponding to that text identifier are integrated into the corresponding hash value. The resulting key-value pairs are then stored in memory to construct a complete hash mapping table structure. This hash mapping table establishes a "key-storage location" correspondence based on a hash function. Each text identifier serves as a unique key, and its corresponding memory storage address can be directly associated with it through hash operations.

[0045] During the construction process, appropriate memory space is allocated based on the total number of index entries to maintain a reasonable table load level and avoid performance degradation due to insufficient space. Because it uses a key-value pair storage format, this structure can achieve fast addressing without relying on pre-sorting of index entries, resulting in greater flexibility in data organization.

[0046] S332B. Based on the hash mapping table structure, establish a direct location path based on hash value calculation. The direct location path is used to directly locate the corresponding hash value by calculating the hash value of the target identifier when the target identifier is received. Specifically, for step S332B, based on the completed hash mapping table structure, a corresponding direct location lookup path is established. When a query request carrying a target identifier is received, there is no need to perform multiple numerical comparisons and interval shrinking operations. The storage location of the corresponding parameter of the target identifier can be directly located by simply calculating the hash value of the target identifier, and the corresponding offset information and content length information can be quickly obtained.

[0047] In practice, a preset hash function is invoked to perform a hash operation on the input target identifier, obtaining the index of the corresponding storage location in the hash mapping table. Then, the memory address corresponding to that location is directly accessed to read the stored hash value data, completing the full location process for the target identifier. Under normal circumstances without hash collisions, this location process requires only one hash calculation and one memory access, and the search time is largely unaffected by the total number of indexes.

[0048] This embodiment constructs a hash mapping table structure with text identifiers as keys and establishes a corresponding hash direct location path to achieve constant time complexity index lookup capability. The lookup speed is not affected by the growth of the index data size, and the location response speed is fast. It can efficiently support high-frequency, high-concurrency text query needs and improve the overall response efficiency of text search.

[0049] Furthermore, in some embodiments, step S4, "mapping the data area to the virtual address space of the current process according to the second storage range, and obtaining the mapped address of the data area in the virtual address space," may specifically include: S41. Determine the starting file offset position of the data area in the text library file and the length of the data to be mapped based on the starting position defined by the second storage range and the total length of the data; Specifically, for step S41, based on the defined second storage range of the data area, the starting position parameter is extracted and used as the starting file offset position for this mapping, which is the byte offset of the first byte of the data area relative to the starting point of the entire text library file; at the same time, the total length of the data corresponding to the second storage range is extracted as the length of the data to be mapped for this mapping. This transforms the broad range into standardized input parameters recognizable by the memory mapping interface, precisely locking the file segment to be mapped, ensuring that mapping is only established for the data area content, avoiding the inclusion of irrelevant areas such as the file header and index area in the mapping range, and preventing invalid occupation of the virtual address space.

[0050] S42. Initiate a memory mapping request to the operating system kernel to which the current process belongs. The memory mapping request includes the starting file offset position and the length of the data to be mapped, in order to request the operating system to map the data area to the virtual address space of the current process. Specifically, in step S42, the memory mapping interface provided by the operating system kernel in the current running environment is invoked to initiate a formal mapping request to the kernel. The request parameters carry the previously determined starting file offset and the length of the data to be mapped, while also specifying mapping access permissions, sharing mode, and other attributes. The request asks the operating system kernel to map the data area of ​​the specified offset and length in the text library file into the virtual address space of the currently running process. This request is uniformly handled and scheduled by the operating system kernel. The mapping process only establishes the correspondence between the file disk data and the virtual memory address; it does not load all the data into physical memory at once. Instead, the kernel completes page loading as needed during subsequent actual access, balancing memory resource consumption and access performance.

[0051] S43. Obtain the virtual memory starting address returned by the operating system kernel, and use the virtual memory starting address as the mapping address of the data area; Specifically, in step S43, after the operating system kernel completes the establishment of the mapping relationship, it returns the starting address of the allocated virtual memory region to the current process, i.e., the virtual memory starting address; this starting address is saved as the mapping address of the data area, serving as the base address for subsequent access to all content in the data area. This address belongs to the private virtual address space of the current process. The process can access content at any location in the data area by adding an offset to this base address. The access method is completely consistent with accessing ordinary memory, and there is no need to call the system interface of the file reading class.

[0052] This embodiment uses a three-step mapping establishment process. First, it accurately locks the file range to be mapped. Then, it completes the address mapping binding through the operating system kernel. Finally, it obtains a directly accessible baseline mapping address. This ensures the accuracy of the mapping range and enables memory-based access to file data by leveraging kernel-level memory management mechanisms. This provides address support for zero-disk I / O reading of text content during subsequent queries, thereby improving the efficiency of reading text content.

[0053] Furthermore, in some embodiments, step S5, "in response to a text query request containing a target identifier, searching for location description information corresponding to the target identifier in the index dataset," may specifically include: S51. Using the target identifier as the search keyword, perform a matching search operation in the index dataset; Specifically, in step S51, the target identifier carried in the query request is extracted, and this target identifier is used as the unique search keyword. A matching search operation is then performed within the index dataset that has been built in memory. The index dataset stores the association between all text identifiers and their corresponding location information. The entire search process is completed in memory without accessing disk storage media. The search process compares the text identifiers in the dataset with the target identifier to determine whether there is a completely matching index entry.

[0054] S52. When a target index information matching the target identifier is found in the index dataset, the corresponding offset information and content length information are extracted from the target index information, and the extracted offset information and content length information are used as location description information. Specifically, for step S52, when the search operation successfully locates the target index information that perfectly matches the target identifier in the index dataset, the offset information and content length information are extracted from the corresponding data fields of the target index information. The offset information indicates the relative starting position of the target text content in the data area, and the content length information indicates the total number of bytes occupied by the target text content. These two pieces of information are integrated as the location description information output by this query and passed to the subsequent text content reading stage.

[0055] S53. When no index information matching the target identifier is found in the index dataset, a search failure indication is generated and the processing flow of the current query request is terminated; Specifically, for step S53, if the search operation completes the retrieval of all valid entries in the index dataset but still fails to find index information matching the target identifier, it indicates that the text content corresponding to the target identifier does not exist in the current text database. At this time, a standardized search failure indication will be generated. This indication can carry information such as the target identifier and a status indicator indicating no matching result, and is used to provide feedback on the query result to the request initiator. After generating the failure indication, the system terminates the subsequent processing flow of the current query request, no longer executes subsequent steps such as reading text content, and directly returns the failure indication.

[0056] This embodiment can accurately obtain the location parameters required for text positioning when the query is successful, providing a precise addressing basis for subsequent text reading. It can also terminate the process and provide status feedback in a timely manner when no matching result is found, making the entire index lookup process logically closed-loop and with clear boundaries, thus ensuring the accuracy and standardization of text query processing.

[0057] Furthermore, in some embodiments, step S6, "reading and returning the target text content from the virtual address space based on the location description information and the mapping address," may specifically include: S61. Obtain the offset information and content length information from the location description information; Specifically, for step S61, the offset information and content length information are extracted from the position description information output by the preceding steps. The offset information represents the relative starting position of the target text content within the entire data area, while the content length information represents the total number of bytes occupied by the target text content. The offset information and content length information are the input basis for subsequent address calculations and data reading; their accuracy directly determines the completeness and correctness of the final read content.

[0058] S62. Calculate the sum of the mapped address and the offset information, and use the calculation result as the starting address for reading the target text content in the virtual address space; Specifically, in step S62, using the mapping base address corresponding to the data area as the reference address, the reference mapping address is summed with the previously extracted offset information. The result is the absolute starting address for reading the target text content in the current process's virtual address space. Since the data area has been fully mapped to the process's virtual address space, the data corresponding to this starting address can be read directly through memory access without needing to call the file system interface to initiate disk I / O operations.

[0059] S63. Allocate an output buffer based on the content length information; Specifically, in step S63, based on the extracted content length information, an output buffer with a matching capacity is allocated in the process memory space. The total number of bytes in the buffer corresponds exactly to the content length of the target text, and it is used to temporarily hold the raw byte data of the target text to be read. The buffer capacity strictly matches the content length, ensuring that all target text data can be fully accommodated while avoiding resource waste caused by allocating redundant memory space.

[0060] S64. Starting from the initial read address, read the raw byte data of the target text content continuously from the virtual address space into the output buffer according to the data length indicated by the content length information, and return the read raw byte data; Specifically, in step S64, starting from the calculated initial read address, the original byte data of the target text content is continuously copied from the virtual address space to the allocated output buffer according to the data length indicated by the content length information. The entire reading process is essentially a memory data copy operation, without triggering any disk I / O requests, and the read speed is on the same order of magnitude as regular memory access. After the data copy is complete, the system returns the original byte data containing the target text data to the request initiator, completing the current text content reading process.

[0061] This embodiment completely transforms text content reading into pure memory-level data operations, eliminating the need to initiate disk file read requests throughout the process. This fundamentally avoids the performance overhead caused by disk I / O, while ensuring the integrity and accuracy of the read data, achieving low-latency and high-efficiency text content reading.

[0062] Furthermore, in some embodiments, the construction method of the index area and data area of ​​the text library file in step S1 includes: S101. Generate an index entry for each text content. The index entry includes a text identifier, offset information, and content length information. The offset information points to the starting storage location of the text content in the data area, and the content length information is used to indicate the storage length of the text content. Specifically, in step S101, all text content to be included in the text library is traversed, and a unique index entry is generated for each individual text content. Each index entry corresponds one-to-one with the text content. Each index entry contains three types of fields: first, a text identifier, which is a unique number for the text content and is used as a search keyword for subsequent queries; second, offset information, which points to the starting storage position of the corresponding text content within the data area, based on the starting position of the data area; and third, content length information, which indicates the total number of storage bytes occupied by the text content. The index entry completely records the positioning parameters of the text content and is the core data carrier for realizing the mapping from text identifier to text content.

[0063] S102. Arrange all index items sequentially according to the size of their text identifiers to form a contiguous index area; Specifically, in step S102, all generated index items are sorted in ascending order according to the numerical value of the text identifiers in each index item, so that all index items are arranged in ascending order of identifiers. After sorting, all index items are concatenated and stored continuously to form a continuous storage area without gaps, which is the index area of ​​the text library. Since each index item uses a fixed-length storage format, the overall structure of the sorted index area is regular, and its total storage length can be directly calculated from the total number of index items and the length of each item, which facilitates subsequent batch reading and boundary positioning.

[0064] S103. Store all text content consecutively after the index area according to the offset information recorded in the index entry, forming the data area; Specifically, in step S103, starting from the end position of the index area, the corresponding text content is stored sequentially and continuously according to the offset information in each index entry. All text content is arranged closely without redundant intervals, forming the data area of ​​the text library. The data area follows the index area, and is stored adjacent to and continuously with the index area, together forming a complete text library file. The actual storage position of each text content corresponds exactly to the offset information recorded in the index entry. The offset information is relative to the starting position of the data area, ensuring that the target text content can be accurately located using the offset.

[0065] This embodiment constructs a regular file structure with index and data partitioned storage, ordered indexes, and compact data by generating one-to-one corresponding index items, building index areas sorted by identifiers, and storing continuously arranged data areas. This ensures the accuracy of the mapping relationship between text identifiers and text content, and provides a suitable file structure foundation for subsequent batch loading of indexes and memory mapping of data areas.

[0066] Furthermore, in some embodiments, before step S33 "parse each index information in the memory buffer", the method further includes: S301. Obtain the total number of index entries recorded in the attribute information, count the actual number of index entries read into the memory buffer, and verify whether the actual number of entries is consistent with the total number of index entries; Specifically, for step S301, firstly, the total number of pre-recorded index entries is obtained from the attribute information of the text library file. This value represents the total number of nominal index entries written during the construction of the text library, indicating the complete entry scale that the index area should contain. Simultaneously, based on the total data length of the memory buffer and the fixed storage length of a single index entry, the actual number of index entries read into the memory buffer is counted. Alternatively, the total number of actual entries can be obtained by iterating through each entry. Then, the actual number of entries is compared with the nominal total number of index entries to verify that they are completely consistent. Anomalies such as incomplete index data loading, partial file corruption, and truncation are investigated to ensure that the number of index entries loaded into memory matches the complete number designed in the file, avoiding omissions in subsequent queries due to missing data.

[0067] S302. Traverse each index information read, obtain the offset information and content length information in each index information, and verify whether the offset information in each index information is less than the total data length of the data area, whether the content length information is greater than zero, and whether the sum of the offset information and the content length information does not exceed the second storage range of the data area. Specifically, for step S302, each index entry is traversed, and its offset and content length information are extracted one by one. Three boundary checks are then performed sequentially: the first check verifies whether the offset value is less than the total length of the data area, ensuring the text content's start position falls within the valid range of the data area; the second check verifies whether the content length value is greater than zero, excluding abnormal index entries with empty content or invalid lengths; the third check verifies whether the sum of the offset and content length information does not exceed the second storage range of the data area, meaning the end position of the text content does not cross the tail boundary of the data area. This ensures the legality and validity of each index entry's location from an address perspective, preventing address out-of-bounds errors and access to illegal memory regions during subsequent text content reading, thus mitigating the risk of read errors and program crashes.

[0068] S303. Traverse each index information read and verify whether the text identifiers in each index information are sequentially increased in order of size; Specifically, in step S303, all index information in the memory buffer is traversed in storage order, and the text identifier values ​​of adjacent indexes are compared sequentially to verify whether the text identifiers of all index information are sequentially increasing in numerical order, ensuring that the entire index sequence maintains a strict ascending order. If the text identifier value of a later index is less than that of the previous one, it is determined that the sorting does not conform to the preset rules.

[0069] This embodiment can identify various problems in advance before index parsing, and ensure the integrity and compliance of the index data loaded into memory from three dimensions: data scale, address boundaries, and structural rules. It avoids possible anomalies in subsequent query processes and improves the operational stability and data reliability of text search methods.

[0070] Furthermore, in some embodiments, after step S43 "obtaining the virtual memory start address returned by the operating system kernel and using the virtual memory start address as the mapping address of the data area", the method further includes: S44. Verify whether the virtual memory starting address returned by the operating system kernel is a non-empty, available memory address; Specifically, in step S44, after obtaining the virtual memory starting address returned by the operating system kernel, the validity of the virtual memory starting address is first verified to determine whether it is a non-empty, usable memory address. Memory mapping requests may fail due to insufficient file access permissions, insufficient process virtual address space, invalid input parameters, etc. In this case, the kernel will return a null pointer or a specific invalid address identifier. If such invalid addresses are used directly to perform subsequent text reading operations, an illegal memory access error will be triggered, causing the program to crash abnormally.

[0071] S45. Obtain the total length of the data area recorded in the attribute information, and verify whether the length of the data to be mapped is consistent with the total length of the data area; Specifically, in step S45, the total length of the pre-recorded data area is extracted from the attribute information of the text library file. This length is the nominal total number of bytes in the data area when the text library was built, representing the actual complete size of the data area. The system then compares this nominal total length with the length of the data to be mapped used when initiating the memory mapping request to verify that the two values ​​are completely consistent. This process checks for mapping range deviations to avoid situations where only a portion of the data area is mapped, or the mapping range exceeds the actual boundary of the data area, due to errors in the initial range calculation or parameter passing. If such deviations exist, subsequent reading of text content with large offsets may result in invalid data being read or triggering memory out-of-bounds access issues.

[0072] This embodiment employs a dual verification mechanism of address validity verification and mapping length consistency verification. This mechanism can identify abnormal issues immediately after memory mapping is established, proactively avoiding potential faults such as illegal memory access and incomplete data reading in subsequent text reading stages, thereby improving the operational reliability of the memory mapping process.

[0073] Furthermore, in some embodiments, step S4, "mapping the data area to the virtual address space of the current process," includes: When initiating a memory mapping request, set the mapping permission to read-only and set the mapping sharing attribute to shared mapping mode; After the mapping relationship is established, in response to text query requests issued by multiple threads in the same process, each thread shares the same mapping address and independently executes the operation of reading the target text content from the virtual address space.

[0074] Specifically, during the process of initiating a memory mapping request to the operating system kernel, two core parameters are configured simultaneously. The first is the mapping permission configuration, which sets the access permission of the mapped region to read-only, restricting the current process to only have read permissions for this mapped virtual memory segment, and not write or modify permissions. Since the data area of ​​the text library is pre-built static data, the entire query process only needs to read the text content and does not need to modify the original data. The read-only permission not only perfectly matches the functional requirements of the business scenario, but also restricts write operations to the mapped region from the underlying access rules, fundamentally avoiding the risk of write data conflicts and accidental data tampering in multi-threaded scenarios. The second is the mapping sharing attribute configuration, which sets the sharing attribute of the mapping to shared mapping mode. That is, this memory mapping segment is globally shared within the current process. The mapped region is kept only once in the process's virtual address space, and all threads within the process can access this mapped region without having to establish a duplicate mapping relationship for each thread.

[0075] After the mapping relationship is successfully established, when multiple threads within the same process initiate text query requests, all threads share the same data area mapping base address, eliminating the need to allocate a dedicated mapping address for each thread. Each thread, upon receiving its query task, independently completes the entire process of address calculation and data reading of the target text: each thread independently calculates the starting read address of the corresponding text content in the virtual address space based on its own query target identifier, and reads the required text data from that address. The entire reading process is independent and does not interfere with each other. Because the mapped area is read-only, all thread operations are pure read operations, eliminating resource contention and data conflicts caused by data writing. Therefore, no additional thread locks are needed for synchronization control, and data inconsistency issues will not occur. Multiple threads can execute text reading operations simultaneously in parallel, and their respective query processes are unaffected by the execution status of other threads.

[0076] This embodiment, through a combination of read-only permissions and shared mapping mode, enables lock-free concurrent reading of the mapped data area by multiple threads within a process while ensuring the security of text data access. This avoids the system resource overhead caused by repeated mapping, eliminates the performance loss caused by thread synchronization, and improves the overall throughput and running efficiency of text search in a multi-threaded environment.

[0077] To facilitate understanding of the text search method based on memory indexing and memory mapping provided in this embodiment, the following example applies the method to the diagnostic text parsing scenario of automotive diagnostic software, providing a more detailed explanation of the text search method based on memory indexing and memory mapping provided in this application. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0078] In a specific embodiment, the text search method based on memory indexing and memory mapping provided in this embodiment has the following specific process: I. Text library file structure design.

[0079] The text library file used in this embodiment is a continuous binary storage format, consisting of three parts: the file header area, the index area, and the data area. The index area and the data area are stored in a continuous manner.

[0080] File header area: Located at the beginning of the file, it is used to record basic information of the text library, including file identifier, version information, index area start address index_start, total number of index entries index_count, length of a single index entry entry_size, data area start address data_start, and checksum bytes, providing parameter basis for subsequent partition reading and mapping.

[0081] The index area consists of multiple fixed-length index entries. Each index entry contains three types of fields: a 4-byte unsigned integer text ID, a 4-byte unsigned integer offset address (the offset relative to the starting address of the data area), and a 4-byte unsigned integer length of the text content. The total length of a single index entry is 12 bytes. All index entries are stored in ascending order of text ID. The sorting is performed using a quicksort algorithm during the text library generation phase, and then the results are written to the file.

[0082] Data area: Stored consecutively following the index area, it contains the raw byte data of all text content. The position and length of each text content are precisely located using the offset and length parameters in the index entry.

[0083] II. System Initialization Phase.

[0084] When the system starts, it performs initialization operations, completes the memory loading of index data and the mapping of data areas, and all disk operations are completed at this stage.

[0085] The index loading and memory index building system first opens the text library file, reads the file header information, and parses it to obtain the starting address and total length of the index area. Depending on the runtime environment, it calls the corresponding file reading interface: the `read` interface for Linux and the `ReadFile` interface for Windows, continuously reading the entire index area from the disk into a pre-allocated memory buffer `index_buffer`. After loading, it parses all index entries in the buffer and can choose two methods to build the memory lookup structure: Ordered array method: The index items are saved as an ordered array structure. The array element corresponds to a structure containing three fields: text_id, offset, and length, maintaining the original file's sorting order by text ID in ascending order; Hash table method: A hash mapping table is constructed using each text ID as the key and the corresponding offset address and content length as the value, supporting queries with O(1) time complexity. After the index structure is constructed, integrity checks are performed: First, it checks whether the actual number of index entries is consistent with the total number of records in the file header; second, it checks whether the text_id of each index entry remains monotonically increasing; third, it checks whether the offset + length of each index exceeds the data area boundary. After the checks pass, the index structure remains resident in memory until the system exits.

[0086] The system determines the starting position and total length of the data area based on the file header information, and calls the memory mapping API provided by the operating system to map the entire data area into the virtual address space of the current process.

[0087] In a Linux environment: First, call `open` to open the text library file, then call `mmap(NULL, data_size, PROT_READ, MAP_SHARED, fd, 0)` to establish the mapping. `PROT_READ` sets read-only permissions, and `MAP_SHARED` sets the shared mapping mode. The starting address of the mapping returned by the call is the base address pointer `base_ptr`. When the system exits, call `munmap(base_ptr, data_size)` to release the mapping.

[0088] In Windows environment: The mapping is established by sequentially calling CreateFile, CreateFileMapping, and MapViewOfFile, with mapping permissions set to shared read mode (FILE_SHARE_READ). The base_ptr and mapping length (data_size) are saved. Upon system exit, UnmapViewOfFile(base_ptr) is called to release the mapping. After mapping is established, validity checks are performed: first, verifying if the returned base_ptr is a valid non-null memory address; second, verifying if the mapping length matches the total length of the data area recorded in the file header. Once the checks pass, the mapping remains in effect.

[0089] III. Text query execution phase.

[0090] After initialization, the system enters a waiting state and continuously listens for text ID query requests; upon receiving a request, the query is completed entirely through memory operations without any disk I / O.

[0091] Upon receiving a query request carrying the target text ID, the memory index lookup performs a lookup based on the memory index structure: If an ordered array structure is used, a binary search is performed: initialize left=0, right=index_count-1, loop to take mid=(left+right) / 2, compare the text_id at the mid position with the target ID; if they are equal, return the corresponding offset and length; if the target value is smaller, update right to mid-1, otherwise update left to mid+1, until the target is found or the interval is invalid.

[0092] If a hash table structure is used, the hash value is calculated based on the target text ID to directly locate the corresponding storage location and obtain the corresponding offset address and content length. Upon successful lookup, the parameters are validated: whether offset is less than the total length of the data area, whether offset + length does not exceed the data area range, and whether length is greater than 0; if validation fails, a data exception error code is returned. If the target text ID is not found, a query failure error is returned.

[0093] After obtaining a valid offset and length from the memory-mapped text, the system first allocates an output buffer: `char* output_buffer = malloc(length + 1)`, allocating an extra byte to store the C string terminator. Then, a memory copy operation is performed: `memcpy(output_buffer, base_ptr + offset, length)`, directly copying the target text content from the virtual address space to the output buffer. Finally, `output_buffer[length] = '\0'` is set, and the complete text content is returned. The caller is responsible for releasing the output buffer after the query is complete. In multi-threaded concurrent query scenarios, because the memory-mapped region is a read-only shared region with no write operations, multiple threads within the same process can share the same mapping base address, independently performing address calculations and data reading without write conflicts, allowing for parallel completion of the query task.

[0094] In a specific embodiment, this embodiment also provides another specific implementation of the text search method based on memory indexing and memory mapping, the specific steps of which are as follows: Initialization phase: Step A: Complete the disk read and memory construction of the index area in one go, and subsequent queries will no longer access the disk of the index area.

[0095] Step B: Establish the association between virtual memory and file data area through the operating system's memory mapping mechanism, without triggering actual I / O.

[0096] Query loop: The system enters a waiting state and continues to listen for query requests.

[0097] Once a request is received, the process immediately enters a pure memory operation flow.

[0098] Query phase: Step C: Complete ID location in memory with a time complexity of O(log n) or O(1).

[0099] Step D: Copy data directly through pointer offset, without fseek or fread, resulting in zero disk I / O.

[0100] In summary, compared with existing technologies, the text search method based on memory indexing and memory mapping provided in this embodiment optimizes the original mode of requiring O(logn) disk I / O operations per search to zero disk I / O during the query phase. Only one index read and one mapping establishment operation are performed during the initialization phase, fundamentally solving the random disk I / O problem caused by frequent fseek operations. Both index lookup and data reading are pure memory operations, resulting in a query speed improvement of more than an order of magnitude, smoothly supporting high-concurrency query scenarios such as batch data stream parsing. The mature operating system memory mapping feature is implemented with clear logic, requiring no complex custom cache scheduling mechanism, making it easy to develop, implement, and maintain, and ensuring high operational stability.

[0101] It should be understood that, although Figure 2 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0102] To facilitate better implementation of the text search method based on memory indexing and memory mapping according to the embodiments of this application, this application also provides a text search device based on memory indexing and memory mapping based on the above-described text search method based on memory indexing and memory mapping. The meanings of the terms used are the same as in the above-described text search method based on memory indexing and memory mapping, and specific implementation details can be found in the description of the method embodiments.

[0103] Please see Figure 3 , Figure 3 The schematic diagram shows the structure of a text search device based on memory indexing and memory mapping provided in this application embodiment. Specifically, this text search device may include a file acquisition module 201, an information determination module 202, an index loading module 203, a mapping establishment module 204, a query response module 205, and a content reading module 206, as follows: The file acquisition module 201 is used to acquire the text library file to be searched. The text library file includes an index area and a data area. The information determination module 202 is used to determine the first storage range of the index area in the text library file and the second storage range of the data area in the text library file based on the attribute information of the text library file. The index loading module 203 is used to read all the index information contained in the index area into memory according to the first storage range, so as to form an index dataset for searching in memory; The mapping establishment module 204 is used to map the data area to the virtual address space of the current process according to the second storage range, and obtain the mapping address of the data area in the virtual address space; The query response module 205 is used to respond to a text query request containing a target identifier and search for location description information corresponding to the target identifier in the index dataset; The content reading module 206 is used to read and return the target text content from the virtual address space based on the location description information and the mapping address.

[0104] Specific limitations regarding the text search device based on memory indexing and memory mapping can be found in the limitations of the text search method based on memory indexing and memory mapping described above, and will not be repeated here. Each module in the aforementioned text search device based on memory indexing and memory mapping can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0105] The text search device based on memory indexing and memory mapping provided in this embodiment loads all index information of the index area into memory to form an index dataset and maps the data area to the process virtual address space, so that text identification query and content reading are both completed based on memory operations. This fundamentally eliminates the performance overhead caused by frequent random disk I / O, reduces the response latency of text search, and thus improves the running efficiency of text content search.

[0106] Furthermore, embodiments of this application also provide an electronic device, such as... Figure 4 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically: The electronic device may include components such as a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, a power supply 303, and an input unit 304. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 301 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 302, and by calling data stored in the memory 302, thereby providing overall monitoring of the electronic device. Optionally, the processor 301 may include one or more processing cores; preferably, the processor 301 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 301.

[0107] The memory 302 can be used to store software programs and modules. The processor 301 executes various functional applications and text search methods based on memory indexing and memory mapping by running the software programs and modules stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 302 may also include a memory controller to provide the processor 301 with access to the memory 302.

[0108] The electronic device also includes a power supply 303 that supplies power to various components. Preferably, the power supply 303 can be logically connected to the processor 301 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 303 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0109] The electronic device may also include an input unit 304, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0110] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 301 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 302 according to the following instructions, and the processor 301 runs the applications stored in the memory 302 to realize various functions, as follows: The process involves: acquiring the text library file to be searched, which includes an index area and a data area; determining the first storage range of the index area and the second storage range of the data area within the text library file based on its attribute information; reading all index information contained in the index area into memory based on the first storage range to form an index dataset for searching in memory; mapping the data area to the virtual address space of the current process based on the second storage range to obtain the mapped address of the data area in the virtual address space; responding to a text query request containing a target identifier, searching for the location description information corresponding to the target identifier in the index dataset; and reading and returning the target text content from the virtual address space based on the location description information and the mapped address.

[0111] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0112] This application embodiment loads all index information of the index area into memory to form an index dataset and maps the data area to the process virtual address space, so that text identification query and content reading are both completed based on memory operations, fundamentally eliminating the performance overhead caused by frequent random disk I / O, reducing the response latency of text search, and thus improving the running efficiency of text content search.

[0113] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0114] Therefore, embodiments of this application provide a storage medium storing multiple instructions that can be loaded by a processor to execute steps in any of the text search methods based on memory indexing and memory mapping provided in embodiments of this application. For example, the instructions can execute the following steps: The process involves: acquiring the text library file to be searched, which includes an index area and a data area; determining the first storage range of the index area and the second storage range of the data area within the text library file based on its attribute information; reading all index information contained in the index area into memory based on the first storage range to form an index dataset for searching in memory; mapping the data area to the virtual address space of the current process based on the second storage range to obtain the mapped address of the data area in the virtual address space; responding to a text query request containing a target identifier, searching for the location description information corresponding to the target identifier in the index dataset; and reading and returning the target text content from the virtual address space based on the location description information and the mapped address.

[0115] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0116] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0117] Since the instructions stored in the storage medium can execute the steps of any of the text search methods based on memory indexing and memory mapping provided in the embodiments of this application, the beneficial effects that any of the text search methods based on memory indexing and memory mapping provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0118] The foregoing has provided a detailed description of a text search method and apparatus based on memory indexing and memory mapping provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A text search method based on memory indexing and memory mapping, characterized in that, include: Obtain the text library file to be searched, the text library file including an index area and a data area; Based on the attribute information of the text library file, determine the first storage range of the index area in the text library file, and determine the second storage range of the data area in the text library file; According to the first storage range, all index information contained in the index area is read into memory to form an index dataset for searching in memory; Based on the second storage range, the data area is mapped to the virtual address space of the current process to obtain the mapped address of the data area in the virtual address space; In response to a text query request containing a target identifier, the location description information corresponding to the target identifier is searched in the index dataset; Based on the location description information and the mapping address, the target text content is read from and returned from the virtual address space.

2. The text search method based on memory indexing and memory mapping according to claim 1, characterized in that, The step of determining the first storage range of the index area in the text library file and the second storage range of the data area in the text library file based on the attribute information of the text library file includes: Read the file header information of the text library file, and extract the starting position of the index area in the text library file and the total number of index items contained in the index area from the file header information. The attribute information includes the file header information. Based on the starting position and the total number of index items, and in conjunction with the preset storage length occupied by a single index item, the ending position of the index area is calculated, and the first storage range is defined by the starting position and the ending position. Obtain the starting position of the data area in the text library file and the total length of the data area, calculate the ending position of the data area based on the starting position and the total length of the data, and define the second storage range using the starting position and the ending position.

3. The text search method based on memory indexing and memory mapping according to claim 1, characterized in that, The step of reading all index information contained in the index area into memory according to the first storage range to form an index dataset for searching in memory includes: Based on the starting position and storage length defined by the first storage range, the file read pointer is positioned at the starting position of the index area; Using the storage length as the total reading volume, all index information contained in the index area is continuously read from the text library file into the memory buffer; The index information in the memory buffer is parsed, and the text identifier in each index information is used as the search key to establish a mapping relationship between each text identifier and the corresponding offset information and content length information, and the mapping relationship is used as the index dataset.

4. The text search method based on memory indexing and memory mapping according to claim 3, characterized in that, The process of establishing a mapping relationship between each text identifier and its corresponding offset and content length information includes: The parsed index information is arranged in order of the size of the text identifier, and an ordered array structure is constructed in memory. Each array element in the ordered array structure includes the text identifier and the offset information and content length information corresponding to the text identifier. Based on the ordered array structure, a binary search path based on text identifier comparison is established. The binary search path is used to locate the array element corresponding to the target identifier by successively splitting the search interval in half when the target identifier is received.

5. The text search method based on memory indexing and memory mapping according to claim 3, characterized in that, The process of establishing a mapping relationship between each text identifier and its corresponding offset and content length information includes: The text identifiers in each parsed index information are used as hash keys, and the corresponding offset information and content length information are used as hash values ​​to construct a hash mapping table structure in memory. Based on the hash mapping table structure, a direct location path is established based on hash value calculation. The direct location path is used to directly locate the corresponding hash value by calculating the hash value of the target identifier when the target identifier is received.

6. The text search method based on memory indexing and memory mapping according to claim 1, characterized in that, The step of mapping the data area to the virtual address space of the current process according to the second storage range, and obtaining the mapped address of the data area in the virtual address space, includes: Based on the starting position and total data length defined by the second storage range, determine the starting file offset position of the data area in the text library file and the data length to be mapped; A memory mapping request is initiated to the operating system kernel to which the current process belongs. The memory mapping request includes the starting file offset position and the length of the data to be mapped, so as to request the operating system to map the data area to the virtual address space of the current process. Obtain the virtual memory starting address returned by the operating system kernel, and use the virtual memory starting address as the mapping address of the data area.

7. The text search method based on memory indexing and memory mapping according to claim 1, characterized in that, The step of responding to a text query request containing a target identifier by searching for location description information corresponding to the target identifier in the index dataset includes: Using the target identifier as the search keyword, perform a matching search operation in the index dataset; When a target index information matching the target identifier is found in the index dataset, the corresponding offset information and content length information are extracted from the target index information, and the extracted offset information and content length information are used as the location description information. When no index information matching the target identifier is found in the index dataset, a search failure indication is generated and the current query request processing flow is terminated.

8. The text search method based on memory indexing and memory mapping according to claim 1, characterized in that, The step of reading and returning target text content from the virtual address space based on the location description information and the mapping address includes: Obtain the offset information and content length information from the location description information; Calculate the sum of the mapped address and the offset information, and use the calculation result as the starting address for reading the target text content in the virtual address space; Allocate an output buffer based on the content length information; Starting from the initial read address, the system continuously reads the raw byte data of the target text content from the virtual address space into the output buffer according to the data length indicated by the content length information, and returns the read raw byte data.

9. The text search method based on memory indexing and memory mapping according to claim 1, characterized in that, The construction methods of the index area and data area of ​​the text library file include: An index entry is generated for each text content. The index entry includes a text identifier, offset information, and content length information, wherein the offset information points to the starting storage position of the text content in the data area, and the content length information is used to indicate the storage length of the text content. All index items are arranged sequentially according to the size of their text identifiers to form a contiguously stored index area; All text content is stored consecutively in the index area according to the offset information recorded in the index entries, forming the data area.

10. A text search device based on memory indexing and memory mapping, characterized in that, include: The file acquisition module is used to acquire the text library file to be searched, the text library file including an index area and a data area; The information determination module is used to determine, based on the attribute information of the text library file, a first storage range of the index area in the text library file and a second storage range of the data area in the text library file; The index loading module is used to read all the index information contained in the index area into memory according to the first storage range, so as to form an index dataset for searching in memory; The mapping establishment module is used to map the data area to the virtual address space of the current process according to the second storage range, and obtain the mapping address of the data area in the virtual address space; The query response module is used to respond to a text query request containing a target identifier and search for location description information corresponding to the target identifier in the index dataset; The content reading module is used to read and return target text content from the virtual address space based on the location description information and the mapping address.