Data processing method and device, computer equipment and computer readable storage medium

By memory mapping the file in the target process and only obtaining the necessary storage space associated data, the problem of excessive memory usage caused by the increase in the number of processes is solved, and efficient storage of data objects and memory saving use is achieved.

CN119987670AActive Publication Date: 2025-05-13NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510087389.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

As the number of processes increases, the physical memory occupied by data file resources will increase linearly, resulting in excessive device memory usage, especially on the server side of multi-process distributed architecture.

Method used

By performing memory mapping processing in the target process, the file is mapped to the virtual address space, only the storage space association data is obtained to generate a data object, and part of the target data in the file is obtained based on the object in response to the access instruction.

Benefits of technology

The data object generated in the target process only contains the associated data with a small amount of data, rather than all the data, thereby reducing the data's memory usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987670A_ABST
    Figure CN119987670A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device, computer equipment and a computer readable storage medium, memory mapping processing is carried out on a first file based on a target process, and a mapping area of the first file in a virtual address space of the target process is obtained, the first file comprises first target data and storage space associated data of the first target data; in response to the first data access instruction, obtaining storage space associated data in the first file based on the mapping area; generating a first data object according to the storage space associated data; based on the first data object and the storage space associated data packaged by the first data object, obtaining part of first target data in the first file; the first data access instruction is responded on the basis of the part of the first target data, the first data object generated in the target process can contain the storage space associated data with the small data size instead of containing all the data of the first file, and occupation of the data on a memory is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing method, apparatus, computer device and computer-readable storage medium. Background Art

[0002] The existing process loading data files are data files that different processes will load separately. Each process loads separately to generate a process-private data file resource, which is stored in the system's physical memory. As the number of processes increases, the physical memory occupied by the data file resources will increase linearly. For devices, especially servers that use a multi-process distributed architecture, there are many processes, and data file resources will occupy a large amount of memory space. Summary of the invention

[0003] The embodiments of the present application provide a data processing method, apparatus, computer device and computer-readable storage medium, which can realize that the first data object generated by the target process contains storage space-associated data with a smaller amount of data, rather than all the data of the first file, thereby reducing the memory occupied by data.

[0004] A data processing method provided in an embodiment of the present application includes:

[0005] Performing memory mapping processing on a first file based on a target process to obtain a mapping area of ​​the first file in a virtual address space of the target process, wherein the first file includes first target data and storage space associated data of the first target data;

[0006] In response to a first data access instruction, acquiring the storage space associated data in the first file based on the mapping area;

[0007] generating a first data object according to the storage space associated data;

[0008] Based on the first data object and the storage space associated data encapsulated by the first data object, obtaining part of the first target data in the first file;

[0009] The first data access instruction is responded to based on the portion of the first target data.

[0010] Accordingly, an embodiment of the present application further provides a data processing device, including:

[0011] A mapping unit, configured to perform memory mapping processing on a first file based on a target process to obtain a mapping area of ​​the first file in a virtual address space of the target process, wherein the first file includes first target data and storage space associated data of the first target data;

[0012] A first acquisition unit, configured to acquire the storage space associated data in the first file based on the mapping area in response to a first data access instruction;

[0013] A generating unit, configured to generate a first data object according to the storage space associated data;

[0014] A second acquisition unit, configured to acquire part of the first target data in the first file based on the first data object and the storage space associated data encapsulated by the first data object;

[0015] A response unit is used to respond to the first data access instruction based on the part of the first target data.

[0016] Correspondingly, an embodiment of the present application also provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute any data processing method provided in the embodiment of the present application.

[0017] Correspondingly, an embodiment of the present application also provides a computer-readable storage medium, which is used to store a computer program, and the computer program is loaded by a processor to execute any data processing method provided in the embodiment of the present application.

[0018] The embodiment of the present application performs memory mapping processing on the first file based on the target process to obtain a mapping area of ​​the first file in the virtual address space of the target process, wherein the first file includes first target data and storage space associated data of the first target data; in response to a first data access instruction, the storage space associated data in the first file is obtained based on the mapping area; a first data object is generated according to the storage space associated data; based on the first data object and the storage space associated data encapsulated by the first data object, part of the first target data in the first file is obtained; and the first data access instruction is responded to based on part of the first target data, so that the first data object generated in the target process includes storage space associated data with a smaller amount of data, rather than all the data of the first file, thereby reducing the memory occupied by data. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 is a flow chart of a data processing method provided in an embodiment of the present application;

[0021] Figure 2 is a corresponding schematic diagram of the data object and the structure provided in the embodiment of the present application;

[0022] Figure 3 It is a schematic diagram of serialization processing provided by an embodiment of the present application;

[0023] FIG4(1) is a UML schematic diagram provided in an embodiment of the present application;

[0024] FIG4(2) is a UML schematic diagram provided in an embodiment of the present application;

[0025] FIG4(3) is a UML schematic diagram provided in an embodiment of the present application;

[0026] Figure 5 is a schematic diagram of a data processing device provided in an embodiment of the present application;

[0027] Figure 6 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0029] The embodiments of the present application provide a data processing method, apparatus, computer equipment and computer-readable storage medium. The data processing apparatus can be integrated in a computer equipment, which can be a server or a terminal.

[0030] The terminal may include a mobile phone, a wearable smart device, a tablet computer, a laptop computer, a personal computer (PC), and a vehicle-mounted computer.

[0031] Among them, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), as well as big data and artificial intelligence platforms.

[0032] It should be noted that the description order of the following embodiments is not intended to limit the preferred order of the embodiments.

[0033] This embodiment will be described from the perspective of a data processing device. The data processing device may be integrated into a computer device, which may be a server or a terminal.

[0034] The present application provides a data processing method, such as Figure 1 As shown, the specific process of the data processing method can be as follows:

[0035] 101. Perform memory mapping processing on a first file based on a target process to obtain a mapping area of ​​the first file in a virtual address space of the target process, wherein the first file includes first target data and storage space associated data of the first target data.

[0036] Among them, memory mapping (MMAP) can map the content of the first file into the virtual address space of the process, the content of the file can be directly accessed as a part of the process memory, the process can directly read and write the content of the first file, and optionally, the first file can be mapped to the virtual address space of the target process by using the mmap system call.

[0037] The first file may include first target data and data related to the storage location of the first target data, i.e., storage space-associated data. The first target data may be, for example, an attribute configuration table of a game, which may include data such as the skills and attributes of a game character, and the effects of game props. The content of the first target data may be different in different application scenarios, and the storage space-associated data may be flexibly configured, for example, including the storage address of the key-value pairs of the first target data, a hash table, etc. It may be specifically set according to the information required for querying the data, and is not limited here.

[0038] There may be one or more target processes, which map the first file to the virtual address space of the target process, so that the target process can obtain storage space-associated data based on the mapping area of ​​the first file in the virtual address space, and then access the target data in the first file based on the storage space-associated data without having to completely load the target data into the target process, thereby reducing the data's occupancy of device memory.

[0039] 102. In response to a first data access instruction, obtain storage space-associated data in a first file based on a mapping area.

[0040] Among them, the first data access instruction can be an access instruction for the first target data in the first file. The target process can access the data in the first file according to the mapping area. The target process accesses the storage space-associated data in the first file according to the first access instruction. Compared with the first target data, the amount of storage space data is very small.

[0041] In one embodiment, the storage space data may include address mapping data, which may include a hash table, a search tree, a linked list, etc. The address mapping data is used to quickly determine the location information of the required data in the storage space associated data according to the data identifier. The step of "obtaining part of the first target data in the first file based on the first data object and the storage space associated data encapsulated by the first data object" includes:

[0042] determining, according to the first data access instruction, a data identifier of data required to respond to the first data access instruction;

[0043] Determining a storage address corresponding to the data identifier based on the first data object and the address mapping data;

[0044] A portion of the first target data is obtained from the first file based on the storage address.

[0045] The data identifier may be an index, a key value, etc. For example, if the first target data is stored in a key-value pair data structure, the data identifier may be the key value of the first target data. The first data access instruction includes the data identifier of the required data.

[0046] In one embodiment, the first data object may encapsulate address mapping data and the storage address of the first target data. The storage address corresponding to the data identifier may be queried based on the address mapping data in the first data object, and the position in the data encapsulated by the first data object may be used to query the corresponding storage address.

[0047] After the storage address is queried, the required data, that is, part of the first target data, can be obtained from the first file.

[0048] In one embodiment, the storage space associated data further includes a starting storage address of the first file and a relative address of the first target data, that is, an address offset of the first target data relative to the storage starting address of the first file. The step of “determining a storage address corresponding to the data identifier based on the first data object and the address mapping data” may include:

[0049] Determine a target address offset corresponding to the data identifier based on the first data object and the address mapping data;

[0050] Based on the start storage address and the target address offset, a storage address of a portion of the first target data is determined.

[0051] The first data object may encapsulate address mapping data and an address offset of the first target data. The target address offset corresponding to the data identifier in the first data object is determined according to the address mapping data encapsulated by the first data object.

[0052] Based on the starting storage address and the address offset, the storage address of part of the first target data is determined. Specifically, the storage address of the first target data may be obtained by adding the starting storage address to the address offset.

[0053] 103. Generate a first data object according to the storage space association data.

[0054] The first data object may be an object in object-oriented programming, such as a python object, etc. The first data object is generated according to the storage space associated data, and the first data object encapsulates the storage space associated data.

[0055] 104. Based on the first data object and storage space association data encapsulated by the first data object, obtain part of the first target data in the first file.

[0056] Based on the first data object and the storage space associated data encapsulated by the first data object, part of the first target data in the first file can be accessed in the first file, and the part of the first target data can be the data required to respond to the first data access instruction.

[0057] 105. Respond to the first data access instruction based on part of the first target data.

[0058] Responding to the first data access instruction based on part of the first target data may be by directly returning part of the first target data, or by performing data operations on part of the first target data to obtain operation results, and then returning the operation results. The data operations on the first target data may include, for example, four arithmetic operations, finding intersections, etc.

[0059] In one embodiment, the first file also includes second target data. The first target data and the second target data can be data of different data types. For example, the first target data can be a tuple, an array, a set, a dictionary, etc., and the second target data can be an integer, a floating point number, a string, etc. The first target data usually contains multiple data. For example, an array contains multiple data, and some of the data may need to be accessed. The second target data is usually accessed as a whole. For the second target data, a second data object can be generated based on the second target data, that is, the data processing method provided by the present application can also include:

[0060] In response to a second data access instruction, obtaining second target data in the first file based on the mapping area;

[0061] A second data object is generated based on the second target data, and a second data access instruction is responded to based on the second data object.

[0062] The specific process of acquiring the second target data may refer to the process of acquiring the storage space associated data of the first target data, which will not be described in detail here.

[0063] After the second data object is generated according to the acquired second target data, a response to the second data access instruction may be made based on the second data object.

[0064] After the first file is memory mapped, the mapping area can be managed. When the mapping area is no longer used, the mapping area is released. That is, in one embodiment, after the step of "memory mapping the first file based on the target process to obtain the mapping area of ​​the first file in the virtual address space of the target process", the data processing method provided by the present application may further include:

[0065] Generate a management object for the mapping area and the first file, where the management object is used to detect the number of data objects that are generated based on the first file and still exist in the target process;

[0066] If it is detected that the number is zero, the recycling mechanism is triggered to destroy the mapping area corresponding to the first file based on the recycling mechanism.

[0067] Among them, the management object can encapsulate a method for managing the mapping area. Through the management object, the number of data objects generated by the target process based on the first file and currently existing can be detected. If the number is not zero, it means that there are still data objects referencing the data of the first file. If the number is zero, it means that there are currently no data objects referencing the data of the first file. When the number is zero, the recycling mechanism is triggered, and the mapping area corresponding to the first file in the target process is destroyed through the recycling mechanism.

[0068] In one embodiment, the first file may be memory mapped by a preset module object, and a preset module object reference management object may be set to ensure that after the data table is loaded, as long as the module object is not actively dereferenced, this part of the mapping area will not be reclaimed, that is, the step of "memory mapping the first file based on the target process" includes:

[0069] Performing memory mapping processing on the first file based on the target process through a preset module object;

[0070] After the step of “generating a management object for the mapping area and the first file”, the data processing method provided in the embodiment of the present application may further include:

[0071] The preset module object is set to reference the management object so that the preset module object triggers a recycling mechanism based on the management object.

[0072] Specifically, a custom object MmapTable is inherited from types.ModuleType, which is responsible for loading, referencing, and reloading the first file, and mapping the data object of the first file into the memory through the mmap system call. In addition, the import hook mechanism is adopted, and the Finder and Loader are customized to realize the module import in the root directory of the first file and return the MmapTable object.

[0073] When the shallow copy function is called to copy the data object, a corresponding copy method may be executed according to the data type of the first object. That is, in one embodiment, after the step of "generating the first data object according to the storage space associated data", the data processing method provided by the embodiment of the present application may further include:

[0074] If the data object is an object of the first data type, in response to calling the shallow copy function for the data object, returning the data object to the shallow copy function;

[0075] If the data object is an object of the second data type, in response to calling the shallow copy function for the data object, a new data object is generated based on the data object, and the new data object and the pointer of the data object point to the same memory address.

[0076] The data object generated in step 103 may be a custom object. For example, for Tuple, Set, and Dict types, the generated custom object may be a custom ShmTuple, ShmSet, or ShmDict type.

[0077] The object of the first data type may be ShmSet or ShmTuple. ShmSet or ShmTuple is regarded as an immutable object, that is, a singleton class. When a shallow copy function (such as the copy function) is called, ShmSet or ShmTuple itself is returned.

[0078] The object of the second data type may be ShmDict. When a shallow copy function (such as the copy function) is called, a new ShmDict is returned, and the elements of the new ShmDict of the original ShmDict all point to the mapping area of ​​the first file.

[0079] When the shallow copy function is called to copy the data object, a corresponding copy method may be executed according to the data type of the first object. That is, in one embodiment, after the step of "generating the first data object according to the storage space associated data", the data processing method provided by the present application may further include:

[0080] If the data object is an object of the third data type, in response to calling a deep copy function for the data object, performing a layer-by-layer copy process on the data object based on the deep copy function to obtain a new data object;

[0081] If the data object is an object of the fourth data type, in response to calling the deep copy function for the data object, the data in the data object is cumulatively copied through the accumulator function to obtain a new data object, so as to perform operations based on the new data object.

[0082] The third data type may include ShmTuple and ShmDict. When a deep copy function (such as deepcopy function) is called, ShmTuple and ShmDict use the relevant implementations of tuple and dict, and implement deep copy by recursively calling the copy method layer by layer.

[0083] The fourth data type may refer to ShmSet. When a deep copy function (such as the deepcopy function) is called, the deep copy is implemented through the reduce function (ie, the accumulator function). The original deepcopy function can implement a deep copy of the ShmSet according to the return value of this method.

[0084] Optionally, the first file in step 101 can be obtained by the following steps:

[0085] Acquire a second file, where the second file contains at least one first target data object;

[0086] Performing data format conversion processing on the data encapsulated by the first target data object to obtain first target data;

[0087] Performing structure conversion processing according to the storage space associated data of the first target data to obtain first structure data;

[0088] A first file is generated based on the first target data and the first structure data.

[0089] The data format of the second file may be different from that of the first file. For example, the second file is a python file and the first file is a binary file. The first file and the second file may also be in other data formats, which are not limited here. In one embodiment, the second file may be a file converted from a table in an Excel or CSV format. For example, the second file may be a python file converted from an Excel table or a CSV table, and the first file may be a binary file obtained by serializing the second file.

[0090] The second file may include at least one first target data object. The data encapsulated by the first target object is converted into a first target data format. The storage space associated data obtained in step 102 may specifically be storage space data corresponding to one of the first target data, so as to generate a data object based on the storage space data to obtain the data encapsulated by the first target data object and convert it into the first target data.

[0091] The storage space associated data of the first target data is acquired to perform structure conversion processing, and the storage space associated data of the first target data is converted into a structure to obtain first structure data.

[0092] Furthermore, data format conversion processing may be performed on the first structured data, and a first file may be generated based on the processed first structured data and the first target data.

[0093] Optionally, the second file also includes a second target data object. The first target data object and the second target data object may be objects of different data types. The first target data object may also be processed before generating the first file. That is, before the step of “generating the first file based on the first target data and the first structure data”, the data processing method provided by the present application may further include:

[0094] Performing structure conversion processing on the data encapsulated by the second target data object to obtain second structure data;

[0095] The step of “generating a first file based on the first target data and the first structure data” may include:

[0096] A first file is generated according to the first target data, the first structure data and the second structure data.

[0097] For example, the data encapsulated by the second target data object may be subjected to structure conversion processing to obtain second structure data; and the first file may be generated according to the first target data, the first structure data, the second structure data, and the structure data of the second type of data object.

[0098] Exemplarily, the second file is a Python file, and for various Python objects in the second file, the corresponding structures are as follows Figure 2 As shown in the figure, for basic types (None, Bool, Int, Float, String, etc.), the structure stores the necessary data for deserialization into Python objects. For container types (Tuple, Set, Dict), the internal structure of the container corresponding to the Python type can be retained, and the pointer is replaced with a relative address relative to the first address of the first file mapping.

[0099] For the container types Tuple, Set, and Dict, custom Python objects (RefShmTuple, RefShmSet, RefShmDict) are used respectively to reference the shared memory data (that is, the first target data in the first file). When accessed, they are generated through the macro definition PyObject_NEW, which only adds a small amount of memory (Python object header data, shared memory object pointer, data start position pointer). The memory occupied by the actual data is shared.

[0100] As can be seen from the above, the embodiment of the present application performs memory mapping processing on the first file based on the target process to obtain a mapping area of ​​the first file in the virtual address space of the target process, wherein the first file includes first target data and storage space associated data of the first target data; in response to a first data access instruction, the storage space associated data in the first file is obtained based on the mapping area; a first data object is generated according to the storage space associated data; based on the first data object and the storage space associated data encapsulated by the first data object, part of the first target data in the first file is obtained; and the first data access instruction is responded to based on part of the first target data, so that the first data object generated in the target process includes storage space associated data with a smaller amount of data, rather than all the data of the first file, thereby reducing the memory occupancy of data.

[0101] In order to further illustrate the data processing method provided by the present application, an example is given below in which the first file is a binary file and the second file is a python file.

[0102] The data processing method provided in the present application may include two parts, one of which is to obtain a first file by serializing the second file, and the other is to perform memory mapping based on the first file and perform deserialization processing to obtain a Python object.

[0103] First, an Excel table or a CSV table can be obtained. The table can be a configuration table in the game or a data table in other application scenarios. The Excel table or the CSV table is converted into a Python file. The Python object in the Python file is converted into a serializable structure, and the structure sequence is written into a binary file to obtain a first file. Exemplarily, the second file and the first file can be as follows Figure 3 As shown, the first file shown in the figure is only a schematic diagram and does not represent the actual stored data content. The dict in the first row of the table is the data type of the python file, and k1, v1, k2, and v2 are the storage addresses of the key values ​​in the table (the storage addresses can be relative addresses).

[0104] The method provided in the present application for serializing the second file to obtain the first file can be output in the form of a Python module. The project only needs to import the module to use it without modifying the logical code of the project.

[0105] In order to ensure that the project script does not need to make additional adjustments after importing the module, and the data operation method is consistent with the operation of Python native types, the present invention mainly performs the following compatibility processing:

[0106] The custom class MmapTable inherits types.ModuleType and is responsible for loading, referencing, and reloading the data table data objects mapped into memory through the mmap system call. In addition, the import hook mechanism is used to customize Finder and Loader to implement the module import in the root directory of the specified configuration table and return the object of the custom type MmapTable.

[0107] Make the return type of the isinstance function correct. That is, when the custom container types ShmTuple, ShmSet, and ShmDict call the isinstance built-in function for the original container types tuple, set, and dict respectively, the return result is still True. In this way, the logic code that originally called this function to determine the instance type can still run correctly.

[0108] Compatible with data serialization modules (such as bson, json, msgpack, etc.) so that custom container data can also be serialized by these modules. Specifically, it can be tried to detect the object type of the data serialization module. For example, ShmDicts inherits collections.abc.Mapping in its implementation, so that the json module can serialize ShmDicts using the serialization scheme of the mapping type. Optionally, serialization can also be performed through a custom encoding method. The data serialization module is used to serialize the first file into data that can be used for network transmission, but is not used to serialize the python file into a binary file.

[0109] The second file is serialized, and Python basic data types (None, Bool, Int, Float, String, etc.) and container types (Tuple, Set, Dict, etc.) in the second file can be processed separately.

[0110] For Python basic data types, type information and numerical information are stored, and the precision of the data can be controlled. For example, Python int type data objects can be directly stored using C's long long type. Although the representation range is reduced, it is sufficient for configuring the table and can reduce the memory usage of data storage.

[0111] For container types, three data structures, ShmTuple, ShmSet, and ShmDict (names for custom data structures in this application), represent binary data of tuple, set, and dict types respectively. When serializing data, pointers to elements in the container are replaced with relative addresses. When obtaining elements, the corresponding data on the shared file is found by adding the address offset to the first address of the data. In the serialization structure design of container objects, the implementation of structures such as hash tables is kept consistent with the Python engine, maintaining the same high performance as the engine code.

[0112] For various Python types in the second file, the corresponding structures are as follows Figure 2 As shown in the figure, for basic types (None, Bool, Int, Float, String, etc.), the structure stores the necessary data for deserialization into Python objects. For container types (Tuple, Set, Dict), the internal structure of the container corresponding to the Python type can be retained, some data related to the deletion and modification operations can be deleted, and the pointer is replaced with a relative address relative to the first address of the first file mapping. List can be converted into tuple Tuple for processing.

[0113] Memory mapping is performed based on the first file and deserialization is performed to obtain a Python object. Specifically, for Python basic type data, the deserialization operation generates a native Python object. For example, the Python int type data object is directly stored using the C long long type, and the deserialization uses the PyLong_FromLongLong function in the Python C API.

[0114] Generating native Python objects is compatible with binary operations of Python basic types, such as addition, subtraction, multiplication, and division. The data of Python basic data types (such as None, True, False, small integers, etc.) are reusable and do not take up additional space. Python's garbage collection mechanism will recycle basic type data immediately when the reference count is 0. Unless the logic is hooked, it will not reside in memory. When basic type data objects are generated, the memory allocation and data copy overhead are small, and the efficiency is controllable.

[0115] For container types, the deserialization operation generates a custom Python object that references the binary data corresponding to the shared file. The deserialization operation does not use tools such as msgpack, pickle, and json, because Python native objects will be generated. Although the first file is shared, the generated Python objects are private to each process, and each generated Python contains all the data in the first file, which is the same as directly importing a Python data table and generating a module object, and cannot achieve the purpose of saving memory.

[0116] UML diagrams of basic data types and container types are as follows Figure 4(1)-Figure 4(3) As shown, ShmObject is the basic type of all classes, and the field type stores an enumeration value indicating what kind of data the class is. The data in the class structure varies depending on the data type. For the basic type, the necessary data for deserialization into a Python object is stored. For the container type, the internal structure of the container of the corresponding Python type is retained, some data related to the deletion and modification operations are deleted, and the pointer is replaced with a relative address relative to the first address of the file mapping. For the container types tuple, set, and dictionary, custom Python types (RefShmTuple, RefShmSet, RefShmDict) are used to reference the shared memory data. When accessing, a custom object is generated through the macro definition PyObject_NEW, which only adds a small amount of memory (Python object header data, shared memory object pointer, data start position pointer), and the memory occupied by the actual data is shared.

[0117] If the first file needs to be updated, a new binary file for sharing can be generated, and the process can remap the new file, or the shared data table can be degraded to the original Python file.

[0118] For mapping space and mapping file, a special Python object (named MmapObj) is generated for management. This object has the address pointer of the mapping space and the handle of the mapping file respectively. When the object is destroyed, the mapping area will be released and the mapping file will be closed. MMapObj is referenced by the MmapTable module object to ensure that after loading the data table, as long as the module object does not actively release the reference, this part of the mapping area will not be recycled. When accessing the container data ShmTuple, ShmSet, ShmDict in the shared file and generating its corresponding reference object, the reference count of MmapObj is +1. Conversely, when the container reference object is destroyed, the reference count of MmapObj is -1. In this way, as long as the shared data is being used, the reference count of MmapObj is always greater than 0, thereby avoiding being destroyed due to triggering Python's garbage collection mechanism, and the corresponding shared memory area will not be released. At this time, the MmapTable module object can load and reference new data, so that new and old data can coexist.

[0119] In summary, it can be achieved that the first data object generated in the target process contains storage space-associated data with a smaller amount of data, rather than all the data of the first file, thereby reducing the memory occupation by data.

[0120] In order to better implement the data processing method provided in the embodiment of the present application, a data processing device is also provided in one embodiment. The meanings of the terms are the same as those in the above data processing method, and the specific implementation details can refer to the description in the method embodiment.

[0121] The data processing device may be integrated into a computer device, such as Figure 5 As shown, the data processing device may include: a mapping unit 301, a first acquisition unit 302, a generation unit 303, a second acquisition unit 304 and a response unit 305, which are specifically as follows:

[0122] (1) A mapping unit 301, configured to perform memory mapping processing on a first file based on a target process to obtain a mapping area of ​​the first file in a virtual address space of the target process, wherein the first file includes first target data and storage space associated data of the first target data.

[0123] (2) A first acquisition unit 302, configured to acquire the storage space associated data in the first file based on the mapping area in response to a first data access instruction.

[0124] (3) A generating unit 303, configured to generate a first data object according to the storage space association data.

[0125] (4) A second acquisition unit 304, configured to acquire a portion of the first target data in the first file based on the first data object and the storage space associated data encapsulated by the first data object.

[0126] In one embodiment, the storage space associated data includes address mapping data, and the second obtaining unit 304 may also be used to:

[0127] Determine, according to the first data access instruction, a data identifier of the data required to respond to the first data access instruction;

[0128] Based on the first data object and the address mapping data, determining a storage address corresponding to the data identifier;

[0129] The part of the first target data is obtained from the first file based on the storage address.

[0130] In one embodiment, the storage space associated data further includes the starting storage address of the first file and the address offset of the first target data. The second acquisition unit 304 may also be used to:

[0131] Based on the first data object and the address mapping data, determining a target address offset corresponding to the data identifier;

[0132] Based on the start storage address and the target address offset, a storage address of the portion of the first target data is determined.

[0133] (5) A response unit 305, configured to respond to the first data access instruction based on the portion of the first target data.

[0134] In one embodiment, the first file further includes second target data, and the data processing device provided by the present application may further include:

[0135] A third acquisition unit, configured to acquire second target data from the first file based on the mapping area in response to a second data access instruction;

[0136] The instruction response unit is used to generate a second data object based on the second target data, and respond to the second data access instruction based on the second data object.

[0137] In one embodiment, after performing memory mapping processing on the first file based on the target process to obtain a mapping area of ​​the first file in the virtual address space of the target process, the data processing method provided by the present application may further include:

[0138] A management object creation unit, used for generating a management object for the mapping area and the first file, wherein the management object is used for detecting the number of data objects generated based on the first file and still existing in the target process;

[0139] The recycling unit is used to trigger a recycling mechanism if it is detected that the above number is zero, so as to destroy the mapping area corresponding to the above first file based on the above recycling mechanism.

[0140] In one embodiment, the mapping unit 301 may also be used for:

[0141] Performing memory mapping processing on the first file based on the target process through a preset module object;

[0142] The data processing device provided in the present application may also include:

[0143] The preset module object is set to reference the management object, so that the preset module object triggers the recycling mechanism based on the management object.

[0144] In one embodiment, the data processing device provided by the present application may further include:

[0145] a first copy unit, configured to return the data object to the shallow copy function in response to calling the shallow copy function for the data object if the data object is an object of the first data type;

[0146] The second copy unit is used to generate a new data object based on the data object in response to calling a shallow copy function for the data object if the data object is an object of the second data type, and the pointer of the new data object and the pointer of the data object point to the same memory address.

[0147] In one embodiment, the data processing device provided by the present application may further include:

[0148] a third copy unit, configured to, if the data object is an object of a third data type, in response to calling a deep copy function for the data object, perform a layer-by-layer copy process on the data object based on the deep copy function to obtain a new data object;

[0149] The fourth copy unit is used for, if the data object is an object of a fourth data type, in response to calling a deep copy function for the data object, to perform cumulative copying of the data in the data object through an accumulator function to obtain a new data object, so as to perform operations based on the new data object.

[0150] In one embodiment, the data processing device provided by the present application may further include:

[0151] A file acquisition unit, used for acquiring a second file, wherein the second file includes at least one first target data object;

[0152] a format conversion unit, configured to perform data format conversion processing on the data encapsulated by the first target data object to obtain the first target data;

[0153] A first conversion unit, configured to perform structure conversion processing according to the storage space associated data of the first target data to obtain first structure data;

[0154] A generating unit is used to generate the first file based on the first target data and the first structure data.

[0155] In one embodiment, the second file further includes a second target data object, and the data processing device provided by the present application may further include:

[0156] A second conversion unit is used to perform structure conversion processing on the data encapsulated by the second target data object to obtain second structure data;

[0157] The above generation unit can be used to:

[0158] The first file is generated based on the first target data, the first structure data and the second structure data.

[0159] As can be seen from the above, the data processing device of the embodiment of the present application performs memory mapping processing on the first file based on the target process through the mapping unit 301, and obtains the mapping area of ​​the first file in the virtual address space of the target process, wherein the first file contains the first target data and the storage space associated data of the first target data; the first acquisition unit 302 responds to the first data access instruction, and obtains the storage space associated data in the first file based on the mapping area; the generation unit 303 generates a first data object according to the storage space associated data; the second acquisition unit 304 obtains part of the first target data in the first file based on the first data object and the storage space associated data encapsulated by the first data object; the response unit 305 responds to the first data access instruction based on the part of the first target data, so that the first data object generated in the target process contains storage space associated data with a smaller amount of data, rather than all the data of the first file, thereby reducing the data occupation of the memory.

[0160] Accordingly, the embodiment of the present application also provides a computer device, which may be a terminal. Figure 6 As shown, Figure 6A schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device 500 includes a processor 501 having one or more processing cores, a memory 502 having one or more computer-readable storage media, and a computer program stored in the memory 502 and executable on the processor. The processor 501 is electrically connected to the memory 502. It will be understood by those skilled in the art that the computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0161] The processor 501 is the control center of the computer device 500. It uses various interfaces and lines to connect the various parts of the entire computer device 500, executes various functions of the computer device 500 and processes data by running or loading software programs and / or modules stored in the memory 502, and calling data stored in the memory 502, thereby monitoring the computer device 500 as a whole.

[0162] In the embodiment of the present application, the processor 501 in the computer device 500 will load instructions corresponding to the processes of one or more application programs into the memory 502 according to the following steps, and the processor 501 will run the application programs stored in the memory 502 to implement various functions:

[0163] Performing memory mapping processing on the first file based on the target process to obtain a mapping area of ​​the first file in the virtual address space of the target process, wherein the first file includes first target data and storage space associated data of the first target data;

[0164] In response to the first data access instruction, acquiring storage space associated data in the first file based on the mapping area;

[0165] Generate a first data object according to the storage space associated data; obtain part of the first target data in the first file based on the first data object and the storage space associated data encapsulated by the first data object;

[0166] A first data access instruction is responded to based on a portion of the first target data.

[0167] As can be seen from the above, the embodiment of the present application performs memory mapping processing on the first file based on the target process to obtain a mapping area of ​​the first file in the virtual address space of the target process, wherein the first file includes first target data and storage space associated data of the first target data; in response to a first data access instruction, the storage space associated data in the first file is obtained based on the mapping area; a first data object is generated according to the storage space associated data; based on the first data object and the storage space associated data encapsulated by the first data object, part of the first target data in the first file is obtained; and the first data access instruction is responded to based on part of the first target data, so that the first data object generated in the target process includes storage space associated data with a smaller amount of data, rather than all the data of the first file, thereby reducing the memory occupancy of data.

[0168] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0169] Optional, such as Figure 6 As shown, the computer device 500 further includes: a touch screen 503, a radio frequency circuit 504, an audio circuit 505, an input unit 506, and a power supply 507. The processor 501 is electrically connected to the touch screen 503, the radio frequency circuit 504, the audio circuit 505, the input unit 506, and the power supply 507, respectively. Those skilled in the art can understand that Figure 6 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0170] The touch display screen 503 can be used to display a graphical user interface and receive operation instructions generated by the user acting on the graphical user interface. The touch display screen 503 may include a display panel and a touch panel. Among them, the display panel may be used to display information input by the user or information provided to the user and various graphical user interfaces of computer equipment, and these graphical user interfaces may be composed of graphics, text, icons, videos and any combination thereof. Optionally, the display panel may be configured in the form of a liquid crystal display (LCD, Liquid Crystal Display), an organic light emitting diode (OLED, Organic Light-Emitting Diode) and the like. The touch panel may be used to collect the user's touch operation on or near it (such as the user using any suitable object or attachment such as a finger, a stylus, etc. on the touch panel or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute corresponding programs. Optionally, the touch panel may include two parts, a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 501, and can receive the command sent by the processor 501 and execute it. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 501 to determine the type of touch event, and then the processor 501 provides a corresponding visual output on the display panel according to the type of touch event. In an embodiment of the present application, the touch panel and the display panel can be integrated into the touch display screen 503 to realize the input and output functions. However, in some embodiments, the touch panel and the touch panel can be used as two independent components to realize the input and output functions. That is, the touch display screen 503 can also be used as a part of the input unit 506 to realize the input function.

[0171] The radio frequency circuit 504 may be used to send and receive radio frequency signals, so as to establish wireless communication with a network device or other computer devices through wireless communication, and to send and receive signals between the network device or other computer devices.

[0172] The audio circuit 505 can be used to provide an audio interface between the user and the computer device through a speaker and a microphone. The audio circuit 505 can transmit the electrical signal converted from the received audio data to the speaker, which is converted into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 505 and converted into audio data, and then the audio data is output to the processor 501 for processing, and then sent to another computer device through the radio frequency circuit 504, or the audio data is output to the memory 502 for further processing. The audio circuit 505 may also include an earphone jack to provide communication between an external headset and the computer device.

[0173] The input unit 506 may be used to receive input numbers, character information or user feature information (such as fingerprint, iris, facial information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0174] The power supply 507 is used to supply power to various components of the computer device 500. Optionally, the power supply 507 can be logically connected to the processor 501 through a power management system, so that the power management system can manage charging, discharging, and power consumption. The power supply 507 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0175] although Figure 6 Not shown, the computer device 500 may also include a camera, a sensor, a wireless fidelity module, a Bluetooth module, etc., which will not be described in detail here.

[0176] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0177] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0178] To this end, an embodiment of the present application provides a computer-readable storage medium, in which multiple computer programs are stored, and the computer program can be loaded by a processor to execute the steps in any data processing method provided in the embodiment of the present application. For example, the computer program can execute the following steps:

[0179] Performing memory mapping processing on the first file based on the target process to obtain a mapping area of ​​the first file in the virtual address space of the target process, wherein the first file includes first target data and storage space associated data of the first target data;

[0180] In response to the first data access instruction, acquiring storage space associated data in the first file based on the mapping area;

[0181] Generate a first data object according to the storage space associated data; obtain part of the first target data in the first file based on the first data object and the storage space associated data encapsulated by the first data object;

[0182] A first data access instruction is responded to based on a portion of the first target data.

[0183] As can be seen from the above, the embodiment of the present application performs memory mapping processing on the first file based on the target process to obtain a mapping area of ​​the first file in the virtual address space of the target process, wherein the first file includes first target data and storage space associated data of the first target data; in response to a first data access instruction, the storage space associated data in the first file is obtained based on the mapping area; a first data object is generated according to the storage space associated data; based on the first data object and the storage space associated data encapsulated by the first data object, part of the first target data in the first file is obtained; and the first data access instruction is responded to based on part of the first target data, so that the first data object generated in the target process includes storage space associated data with a smaller amount of data, rather than all the data of the first file, thereby reducing the memory occupancy of data.

[0184] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0185] The storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0186] The above is a detailed introduction to a data processing method, device, computer equipment and computer storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A data processing method, characterized in that: include: Performing memory mapping processing on a first file based on a target process to obtain a mapping area of ​​the first file in a virtual address space of the target process, wherein the first file includes first target data and storage space associated data of the first target data; In response to a first data access instruction, acquiring the storage space associated data in the first file based on the mapping area; generating a first data object according to the storage space associated data; Based on the first data object and the storage space associated data encapsulated by the first data object, obtaining part of the first target data in the first file; The first data access instruction is responded to based on the portion of the first target data.

2. The method according to claim 1, characterized in that The storage space associated data includes address mapping data, and acquiring part of the first target data in the first file based on the first data object and the storage space associated data encapsulated by the first data object includes: determining, according to the first data access instruction, a data identifier of data required to respond to the first data access instruction; Determining a storage address corresponding to the data identifier based on the first data object and the address mapping data; The portion of first target data is obtained from the first file based on the storage address.

3. The method according to claim 2, characterized in that The storage space associated data also includes a starting storage address of the first file and an address offset of the first target data. The determining of the storage address corresponding to the data identifier based on the first data object and the address mapping data includes: Determining a target address offset corresponding to the data identifier based on the first data object and the address mapping data; Based on the start storage address and the target address offset, a storage address of the portion of the first target data is determined.

4. The method according to claim 1, characterized in that: The first file also includes second target data, and the method further includes: In response to a second data access instruction, acquiring the second target data in the first file based on the mapping area; A second data object is generated based on the second target data, and the second data access instruction is responded to based on the second data object.

5. The method according to claim 1, characterized in that After performing memory mapping processing on the first file based on the target process to obtain a mapping area of ​​the first file in the virtual address space of the target process, the method further includes: Generate a management object for the mapping area and the first file, the management object being used to detect the number of data objects generated based on the first file and still existing in the target process; If it is detected that the number is zero, a recycling mechanism is triggered to destroy the mapping area corresponding to the first file based on the recycling mechanism.

6. The method according to claim 5, characterized in that The memory mapping process of the first file based on the target process includes: Performing memory mapping processing on the first file based on the target process through a preset module object; After generating the management object for the mapping area and the first file, the method further includes: The preset module object is set to reference the management object, so that the preset module object triggers the recycling mechanism based on the management object.

7. The method according to claim 1, characterized in that After generating the first data object according to the storage space associated data, the method further includes: If the data object is an object of the first data type, in response to calling a shallow copy function for the data object, returning the data object to the shallow copy function; If the data object is an object of the second data type, in response to calling a shallow copy function for the data object, a new data object is generated based on the data object, and the pointer of the new data object and the pointer of the data object point to the same memory address.

8. The method according to claim 1, characterized in that After generating the first data object according to the storage space associated data, the method further includes: If the data object is an object of the third data type, in response to calling a deep copy function for the data object, performing a layer-by-layer copy process on the data object based on the deep copy function to obtain a new data object; If the data object is an object of the fourth data type, in response to calling a deep copy function for the data object, the data in the data object is cumulatively copied through an accumulator function to obtain a new data object, so as to perform operations based on the new data object.

9. The method according to any one of claims 1 to 8, characterized in that: The first file is obtained by the following steps: Acquire a second file, wherein the second file includes at least one first target data object; Performing data format conversion processing on the data encapsulated by the first target data object to obtain the first target data; Performing structure conversion processing according to the storage space associated data of the first target data to obtain first structure data; The first file is generated based on the first target data and the first structure data.

10. The method according to claim 9, characterized in that The second file further includes a second target data object. Before generating the first file based on the first target data and the first structure data, the method further includes: Performing structure conversion processing on the data encapsulated by the second target data object to obtain second structure data; The generating the first file based on the first target data and the first structure data includes: The first file is generated according to the first target data, the first structure data and the second structure data.

11. A data processing device, characterized in that: include: A mapping unit, configured to perform memory mapping processing on a first file based on a target process to obtain a mapping area of ​​the first file in a virtual address space of the target process, wherein the first file includes first target data and storage space associated data of the first target data; A first acquisition unit, configured to acquire the storage space associated data in the first file based on the mapping area in response to a first data access instruction; A generating unit, configured to generate a first data object according to the storage space associated data; A second acquisition unit, configured to acquire part of the first target data in the first file based on the first data object and the storage space associated data encapsulated by the first data object; A response unit is used to respond to the first data access instruction based on the part of the first target data.

12. A computer device, characterized in that: It comprises a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the data processing method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and the computer program is loaded by a processor to execute the data processing method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • File sharing method and device, electronic equipment and storage medium

    CN115934662A

  • Data processing method and device, storage medium and electronic equipment

    CN116028455A

  • Data processing method and device, computer readable storage medium and computer equipment

    CN119149261A

  • Data Object Profiling During Program Execution

    US20130067192A1

  • File access method and apparatus, and storage system

    US20170168952A1