Instruction translation method and device, electronic equipment and storage medium
By using tree data structure and symbol information reconstruction in the symbol table file, the mapping relationship between target address, function number and line number is established, and the problem of low instruction translation efficiency in the existing technology is solved, and a more efficient instruction translation process is realized.
Patent Information
- Application Number
- CN202510134487.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-16
AI Technical Summary
When the client crashes and stutters occur in the existing technology, the instructions are translated by reading symbol table files, which has a high time complexity and the translation efficiency needs to be improved.
The symbol table file is reconstructed using tree data structure and symbol information, and the mapping relationship between target address, function number and line number is established based on tree data structure and symbol information, and the retrieval process of symbol dictionary is optimized.
It reduces the time complexity of instruction translation, improves the efficiency of instruction translation, and reduces the size of symbol table files.
Smart Images

Figure CN120010920A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to an instruction translation method, device, electronic device and storage medium. Background Art
[0002] At present, when the client crashes and freezes, it usually reads the symbol table file and obtains the function character and line number based on the instruction translation request mapping to facilitate the location of the fault. In related technologies, the binary search method is generally used to search for function characters and line numbers in the symbol table file, which has a high time complexity and the translation efficiency needs to be improved. Summary of the invention
[0003] The following is a summary of the subject matter of the detailed description of the present disclosure. This summary is not intended to limit the scope of the claims.
[0004] The embodiments of the present disclosure provide an instruction translation method, device, electronic device and storage medium, which can reduce the time complexity of instruction translation and improve the efficiency of instruction translation.
[0005] In one aspect, an embodiment of the present disclosure provides an instruction translation method, comprising:
[0006] receiving an instruction translation request, and when the instruction translation request carries a target address of an original instruction, obtaining the target address from the instruction translation request;
[0007] Obtaining a first symbol table file, retrieving a target node in a tree data structure of the first symbol table file according to the target address, and reading a symbol offset from the target node;
[0008] Retrieving the function number of the target address and the first row number where the original instruction is located in the first symbol information of the first symbol table file according to the symbol offset, wherein the first symbol information is used to store a mapping relationship between the symbol offset, the function number and the first row number;
[0009] A first symbol dictionary is obtained according to the target address, a first function character of the original instruction is retrieved from the first symbol dictionary according to the function number, and a translation result of the original instruction is obtained based on the first function character and the first line number.
[0010] On the other hand, the embodiment of the present disclosure further provides an instruction translation device, including:
[0011] A request processing module, configured to receive an instruction translation request, and when the instruction translation request carries a target address of an original instruction, obtain the target address from the instruction translation request;
[0012] A tree data structure processing module, used for obtaining a first symbol table file, retrieving a target node in the tree data structure of the first symbol table file according to the target address, and reading a symbol offset from the target node;
[0013] A symbol information processing module, used for retrieving the function number of the target address and the first line number of the original instruction from the first symbol information of the first symbol table file according to the symbol offset, wherein the first symbol information is used for storing a mapping relationship between the symbol offset, the function number and the first line number;
[0014] The symbol dictionary processing module is used to obtain a first symbol dictionary according to the target address, retrieve a first function character of the original instruction in the first symbol dictionary, and obtain a translation result of the original instruction based on the first function character and the first line number.
[0015] Optionally, the tree data structure processing module is further used for:
[0016] Retrieving a first reference address matching the target address from a root node of the tree data structure of the first symbol table file, wherein the first reference address is a starting address of an address range where the target address is located;
[0017] Retrieving a second reference address matching the target address from a branch node associated with the first reference address;
[0018] The leaf node associated with the second reference address is determined as the target node.
[0019] Optionally, the tree data structure processing module is further used for:
[0020] Determine an element position of the first reference address in the root node, and determine a node address offset according to the element position;
[0021] Obtaining a node start address of a leaf node associated with the root node, and determining a destination node address according to the node start address and the node address offset;
[0022] A second reference address matching the target address is retrieved from a branch node corresponding to the destination node address.
[0023] Optionally, the tree data structure processing module is further used for:
[0024] Obtaining, in the instruction translation request, a system identifier of an operating system from which the instruction translation request comes, and determining a memory block size corresponding to the operating system according to the system identifier;
[0025] The node address offset is determined according to the product of the element position and the memory block size.
[0026] Optionally, the symbol dictionary processing module is further used to:
[0027] Obtain a tag file, and retrieve a symbol dictionary index in the tag file according to the target address, wherein the symbol dictionary index is used to indicate an address range where the target address is located;
[0028] A first symbol dictionary is obtained from a plurality of first candidate dictionaries stored in the shards according to the first symbol table file identifier and the symbol dictionary index.
[0029] Optionally, the symbol dictionary processing module is further used to:
[0030] Obtaining a service object identifier of each service object, a first original dictionary of each service object, and a first symbol table file identifier;
[0031] Divide the first original dictionary into a plurality of first candidate dictionaries, and assign a corresponding symbol dictionary index to each of the first candidate dictionaries;
[0032] The tag file, the first symbol table file identifier, the service object identifier, a plurality of the symbol dictionary indexes, and a plurality of the first candidate dictionaries are stored in association.
[0033] Optionally, the tree data structure processing module is further used for:
[0034] Obtaining a first symbol table file in a local disk according to the first symbol table file identifier;
[0035] When the first symbol table file does not exist in the local disk, obtaining the first symbol table file in the network file system according to the first symbol table file identifier;
[0036] When the first symbol table file does not exist in the network file system, the first symbol table file is obtained in the cloud object storage according to the first symbol table file identifier.
[0037] Optionally, the tree data structure processing module is further used for:
[0038] Determining the usage popularity of the first symbol table file;
[0039] The storage location of the first symbol table file is adjusted according to the usage heat, wherein the storage location includes the local disk, the network file system and the cloud object storage, and the usage heat corresponding to the local disk, the network file system and the cloud object storage decreases in sequence.
[0040] Optionally, the instruction translation device further includes a symbol table file making module, and the symbol table file making module is used to:
[0041] Obtain an original symbol table file, and extract from the original symbol table file a plurality of the target addresses, the first function characters corresponding to the respective target addresses, and the first row numbers corresponding to the respective first function characters;
[0042] Creating the tree data structure based on the plurality of target addresses and the symbol offsets, assigning the function number to the first function character, creating the first symbol information based on the symbol offset, the function number and the first row number, and encapsulating the tree data structure and the first symbol information into the first symbol table file;
[0043] The first symbol dictionary is created based on the function number and the first function character.
[0044] Optionally,
[0045] The request processing module is also used for obtaining the instruction to be translated and the second symbol table file identifier in the instruction translation request when the instruction translation request carries the instruction to be translated, wherein the instruction to be translated includes the characters to be translated and the original instruction characters in the original instruction except the second function characters, and the characters to be translated are the characters before translation corresponding to the second function characters;
[0046] The tree data structure processing module is further used to obtain a second symbol table file according to the second symbol table file identifier, and retrieve the second row number of the character to be translated in the second symbol dictionary according to the second symbol information of the character to be translated, wherein the second symbol information is used to store the mapping relationship between the character to be translated and the second row number;
[0047] The symbol dictionary processing module is also used to obtain the second symbol dictionary according to the character to be translated, and obtain the second function character and the third row number where the original instruction is located in the second symbol dictionary according to the second row number as the translation result of the instruction to be translated.
[0048] Optionally, the symbol dictionary processing module is further used to:
[0049] Determine a plurality of second candidate dictionaries stored in the fragment according to the second symbol table file identifier, and match the characters to be translated with instruction indexes corresponding to the respective second candidate dictionaries, wherein the instruction indexes include index characters;
[0050] When there is no instruction index matching the character to be translated, sorting the index characters in a preset order to obtain an index character sequence;
[0051] Determine a first character and a second character adjacent to the character to be translated in the index character sequence, wherein the ranking of the first character in the index character sequence is lower than the ranking of the second character in the index character sequence;
[0052] The second candidate dictionary corresponding to the instruction index where the first character is located is obtained as the second symbol dictionary.
[0053] Optionally, the symbol dictionary processing module is further used to:
[0054] Obtaining a service object identifier of each service object, a second original dictionary of each service object, and a second symbol table file identifier, and obtaining a plurality of second function characters and the second row number corresponding to each second function character in the second original dictionary;
[0055] creating a second candidate dictionary, sequentially writing the second function characters and the corresponding second row numbers as entries into the second candidate dictionary until the storage space occupied by the second candidate dictionary reaches a preset threshold, creating the next second candidate dictionary to continue writing the entries until all the second function characters and the corresponding second row numbers are written, and assigning the corresponding instruction index to each second candidate dictionary;
[0056] The second symbol table file identifier, the service object identifier, a plurality of the instruction indexes and a plurality of the second candidate dictionaries are stored in association with each other.
[0057] On the other hand, an embodiment of the present disclosure further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned instruction translation method when executing the computer program.
[0058] On the other hand, an embodiment of the present disclosure further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is executed by a processor to implement the above-mentioned instruction translation method.
[0059] On the other hand, the embodiment of the present disclosure further provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above-mentioned instruction translation method.
[0060] The disclosed embodiment includes at least the following beneficial effects: by receiving an instruction translation request, when the instruction translation request carries the target address of the original instruction, obtaining the target address in the instruction translation request, obtaining the first symbol table file, retrieving the target node in the tree data structure of the first symbol table file according to the target address, reading the symbol offset from the target node, retrieving the function number of the target address and the first line number of the original instruction from the first symbol information of the first symbol table file according to the symbol offset, because the first symbol table file contains the tree data structure and the first symbol information, compared with the symbol table file in the related art, the content is reconstructed, and the target address, function number and first line number can be mapped based on the tree data structure and the first symbol information, so as to achieve file size reduction, and on this basis, obtaining the first symbol dictionary according to the target address, retrieving the first function character of the original instruction in the first symbol dictionary according to the function number, and obtaining the translation result of the original instruction based on the first function character and the first line number, the above translation process optimizes the retrieval process compared with the binary search method, can reduce the time complexity of instruction translation, and improve instruction translation efficiency.
[0061] Other features and advantages of the present disclosure will be set forth in the following description, and in part will be apparent from the description, or may be learned by practicing the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The accompanying drawings are used to provide further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation on the technical solution of the present disclosure.
[0063] Figure 1 A schematic diagram of an optional implementation environment provided for an embodiment of the present disclosure;
[0064] Figure 2 An optional flowchart of the instruction translation method provided in the embodiment of the present disclosure;
[0065] Figure 3 A schematic diagram of the structure of a tree data structure in a symbol table file provided in an embodiment of the present disclosure;
[0066] Figure 4 A flowchart of the instruction translation service provided by the embodiment of the present disclosure;
[0067] Figure 5 A schematic diagram of the structure of a symbol table file and a symbol dictionary proposed in an embodiment of the present disclosure;
[0068] Figure 6 A schematic diagram of an indexing process of a tree data structure provided by an embodiment of the present disclosure;
[0069] Figure 7 A schematic diagram of the effect of associating and storing a symbol dictionary and a symbol table file provided in an embodiment of the present disclosure;
[0070] Figure 8 A schematic diagram of a symbol dictionary indexing process provided by an embodiment of the present disclosure;
[0071] Fig. 9 A schematic diagram of a symbol dictionary indexing process provided by another embodiment of the present disclosure;
[0072] Fig.10 A schematic diagram of a first symbol table file reading process provided by an embodiment of the present disclosure;
[0073] Fig.11 A schematic diagram of the effect of data hierarchical caching provided by an embodiment of the present disclosure;
[0074] Fig.12 A schematic diagram of a process for obtaining a second symbol dictionary provided in an embodiment of the present disclosure;
[0075] Fig.13 A schematic diagram of an instruction translation method in the related art;
[0076] Fig.14 A schematic diagram of the effect of instruction translation provided by an embodiment of the present disclosure;
[0077] Fig.15 An optional specific flow chart of the instruction translation method provided in the embodiment of the present disclosure;
[0078] Fig.16 An optional structural diagram of an instruction translation device provided in an embodiment of the present disclosure;
[0079] Fig.17 A partial structural block diagram of a terminal provided in an embodiment of the present disclosure;
[0080] Fig.18 A partial structural block diagram of a server provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0081] In order to make the purpose, technical solution and advantages of the present disclosure more clear, the present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.
[0082] It should be noted that in various specific embodiments of the present disclosure, when it comes to the need to perform relevant processing based on data related to the characteristics of the target object such as the target object attribute information or attribute information set, the permission or consent of the target object will be obtained first, and the collection, use and processing of these data will comply with relevant laws, regulations and standards. Among them, the target object can be a user. In addition, when the embodiment of the present disclosure needs to obtain the attribute information of the target object, the separate permission or separate consent of the target object will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary target object-related data used to enable the normal operation of the embodiment of the present disclosure will be obtained.
[0083] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.
[0084] To facilitate understanding of the technical solution provided by the embodiments of the present disclosure, some key terms used in the embodiments of the present disclosure are explained here:
[0085] The symbol table file is a data structure generated by the compiler when compiling the code. It is used to store various names in the original instructions (such as functions, variables, classes, etc.) and their specific locations (memory addresses) when the program is running, as well as other details (such as type, scope, line number, etc.). The symbol table file is a key tool for the compiler or interpreter to analyze, optimize, generate and debug the code.
[0086] Cloud Object Storage (COS) is a distributed storage service for storing massive files provided by cloud platforms. Users can store and view data at any time through the network. Object storage is a non-hierarchical data storage method in cloud storage. It does not use a directory tree structure. Each individual data (object) unit exists at the same level in the storage pool. Each object has a unique identification name that can be used for listing and retrieval operations. In addition, each object can also contain metadata.
[0087] The Network File System (NFS) is a system that allows users to access and operate files and directories on remote servers through the network. It achieves access and operation of remote files and directories by mounting the remote file system on the local computer. When a local client requests access to a remote file or directory, NFS passes the request to the remote server, and the server returns the required data or information. In this way, different computer systems can share and access the same data.
[0088] At present, when using the client of the application, freezes and crashes are often encountered. When the client crashes and freezes, the problem can be analyzed and located through the running log. Usually, the running log is not intuitive and requires developers to analyze it in conjunction with the symbol table file. Generally, by reading the symbol table file, the function character and line number are obtained based on the instruction translation request mapping, which is convenient for locating the fault location. In the related technology, the binary search method is generally used to search for function characters and line numbers in the symbol table file, which has a high time complexity and the translation efficiency needs to be improved.
[0089] Based on this, the embodiments of the present disclosure propose an instruction translation method, device, electronic device and storage medium. By reconstructing the content of the symbol table file, the symbol table file contains a tree data structure and symbol information. The target address, function number and line number can be mapped based on the tree data structure and symbol information, thereby reducing the size of the symbol table file and accelerating file reading. On this basis, the target address is used to obtain the symbol dictionary, and the first function character of the original instruction is retrieved from the symbol dictionary according to the function number. The translation result of the original instruction is obtained based on the function character and the line number. Compared with the binary search method, the above translation process optimizes the retrieval process, can reduce the time complexity of instruction translation, and improve the efficiency of instruction translation.
[0090] The method provided in the embodiments of the present disclosure can be applied to different scenarios, including but not limited to cloud technology, artificial intelligence, Internet services, smart transportation, assisted driving and the like.
[0091] Reference Figure 1 , Figure 1 A schematic diagram of an optional implementation environment provided for an embodiment of the present disclosure, the implementation environment includes a terminal 101 and a server 102, wherein the terminal 101 and the server 102 are connected via a communication network.
[0092] The terminal 101 may be a mobile phone, a computer, an intelligent voice interaction device, an intelligent wearable device, an intelligent home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 101 and the server 102 may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present disclosure.
[0093] Exemplarily, terminal 101 can collect crash address information, module information (including symbol table file identifier), device information, application status, crash type and other data of the original instruction run when a client program crashes, convert the collected original data into a specific format through a software development kit (SDK), generate an operation log, and generate an instruction translation request based on the operation log in order to map the crash address information of the operation log into readable stack information, and send the instruction translation request to the server.
[0094] Server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In addition, server 102 can also be a node server in a blockchain network.
[0095] Exemplarily, the server 102 can receive an instruction translation request sent by the terminal 101 and parse the instruction translation request. When the parsed data contains the crash address information and the symbol table file identifier of the original instruction, the crash address information is used as the target address, and the corresponding symbol table file is obtained using the symbol table file identifier. The target node is retrieved from the tree data structure of the symbol table file according to the target address, and the symbol offset is read from the target node. Since the symbol table file can store the mapping relationship between the symbol offset, the function number and the line number, the function number of the target address and the line number of the original instruction can be retrieved from the symbol information of the symbol table file according to the symbol offset. Then, the symbol dictionary can be obtained according to the target address, the function character of the original instruction can be retrieved from the symbol dictionary according to the function number, and the translation result of the original instruction can be obtained based on the function character and the line number.
[0096] like Figure 1As shown, in an application scenario, assuming that the terminal 101 is a computer, the terminal 101 can run an application, the application can be sent to the terminal 101 by the server 102, or can be pre-installed in the terminal 101, and the application can include a pre-configured software development kit, the software development kit can be pre-configured by the server 102 and sent to the terminal 101, and run by the terminal 101, so when the terminal 101 runs the application and crashes or freezes, the software development kit can collect the target address of the original instruction run when the application crashes or freezes, and the corresponding first symbol table file mark The server 102 receives the instruction translation request and parses the instruction translation request. When the instruction translation request carries the target address of the original instruction, the server 102 extracts the target address and the first symbol table file identifier from the instruction translation request. The server 102 then searches and calls the first symbol table file stored in the server 102 or in a database connected to the server 102 according to the first symbol table file identifier. The server 102 retrieves the target address from the tree data structure of the first symbol table file according to the target address. Node, read the symbol offset from the target node, use the first symbol information in the first symbol table file to store the mapping relationship between the symbol offset, the function number and the first row number, retrieve the function number of the target address and the first row number where the original instruction is located in the first symbol information of the first symbol table file according to the symbol offset, then, according to the target address, search and call the first symbol dictionary stored in the server 102 or in the database connected to the server 102, obtain the translation result of the original instruction in the first row number according to the function number, and then return the translation result of the original instruction to the terminal 101 for the user to locate and repair the problem of the application. It can be seen that the symbol table file proposed in the embodiment of the present disclosure has been reconstructed, and the symbol table file contains a tree data structure and the first symbol information, so that the target address, the function number and the first row number can be mapped based on the tree data structure and the first symbol information, so as to achieve the reduction of the file size. On this basis, the first symbol dictionary is obtained according to the target address, and the first function character is retrieved using the first symbol dictionary, and then the translation result of the original instruction is obtained by the first function character and the first row number. In this way, the retrieval stage in the instruction translation process is optimized, the time complexity of instruction translation is effectively reduced, and the instruction translation efficiency is improved.
[0097] Reference Figure 2 , Figure 2An optional flow chart of an instruction translation method provided for an embodiment of the present disclosure, the instruction translation method can be executed by a server, or can also be executed by a terminal, or can also be executed by a server in cooperation with a terminal, the instruction translation method includes but is not limited to the following steps 201 to 204.
[0098] Step 201: receiving an instruction translation request, and when the instruction translation request carries a target address of an original instruction, obtaining the target address from the instruction translation request.
[0099] Among them, the instruction translation request may refer to the data collected and reported by the client SDK when the application crashes or an exception occurs, in the expectation that the server will match the corresponding symbol table file based on this data information. The original instruction may refer to the code instruction executed in the application, which is one of the instructions in all the running codes of the application. When the application crashes or an exception occurs while executing a certain code instruction, the address of the corresponding original instruction is converted into readable symbolic information (such as class name, function name, line number). It is worth noting that the client SDK can obtain the crash context through signal processing or exception capture, extract the target address of the original instruction from the crash context, and parse the module to which the target address belongs (such as the main program, dynamic library), thereby obtaining the symbol table file identifier, loading base address, version information and other original data corresponding to the module, and then assemble these original data into an instruction translation request according to a specific data structure. When the symbol translation service node, i.e., the server, receives the instruction translation request, it can parse the instruction translation request, directly extract the target address of the original instruction and the first symbol table file identifier from the instruction translation request, or extract the field information associated with the target address and the first symbol table file identifier from the instruction translation request, process these associated field information, obtain the target address and the first symbol table file identifier, and then translate the original instruction based on the target address and the first symbol table file identifier response request.
[0100] The target address may refer to the machine instruction address that causes a crash or exception when the application is running. When the application crashes or an exception occurs, the program counter (PC) will point to the address of the currently executed instruction. This instruction address is a memory address that can correspond to an offset in a binary file. The target address needs to be translated into a developer-readable class name, function name, and line number so that the developer can locate and fix the problem. The client SDK can obtain the target address through a signal handler or exception handler, and fill the target address into the instruction translation request, so that the target address can be extracted from the instruction translation request to the address-related field during the translation process.
[0101] Among them, the symbol table file identifier can refer to the unique identifier of the symbol table file used to translate the target address. Each binary file, such as an executable file and a dynamic link library, generates a corresponding symbol table file during compilation. The symbol table file records the mapping relationship between the runtime address and the class name, function name, and parameter return value, etc. Different operating systems and platforms will have symbol table files of different formats. For example, the iOS platform uses the dSYM file in the DWARF format, the Windows platform uses the binary symbol table file in the PDB format, the Linux system uses the so file in the DWARF format, and the Android platform uses the symbol table file in the text format. However, symbol table files of different formats have a unique identifier, namely the symbol table file identifier. The client SDK can extract the symbol table file identifier by locating the corresponding fields of each symbol table file and fill it into the instruction translation request, so that the required symbol table file identifier can be extracted from the instruction translation request by locating the corresponding fields during the translation process.
[0102] Step 202: Obtain a first symbol table file, retrieve a target node in the tree data structure of the first symbol table file according to the target address, and read a symbol offset from the target node.
[0103] Among them, before executing the symbol translation process, it is necessary to load the required symbol table file first. The symbol table file can be stored in different locations, for example, in a local cache, a shared file system, a cloud object storage, a distributed file system, etc. Since the symbol table file corresponds one-to-one to the symbol table file identifier, the server or terminal can construct a storage path through the first symbol table file identifier to achieve accurate positioning of the first symbol table file, and then load and read the corresponding first symbol table file. The first symbol table file identifier can be carried through the instruction translation request.
[0104] It is worth noting that the symbol table file proposed in the embodiment of the present disclosure has been reconstructed through content reconstruction, and the index file and data file of the symbol table file have been trimmed and optimized to streamline the content of the symbol table file and compress the size of the symbol table file, thereby improving the loading and reading speed of the symbol table file and improving the efficiency of data traversal on the symbol table file. Specifically, the data file of the first symbol table file stores symbol information corresponding to each address, and the index file of the first symbol table file is a tree data structure. The tree data structure can be used to retrieve the symbol information corresponding to the target address in the data file. The address ranges of multiple symbols (i.e., class names, function names) are stored in the tree data structure. The address ranges of each symbol are divided and stored in each node in the tree data structure respectively. The tree data structure includes a root node, a branch node, and a leaf node. The branch node can refer to a child node of the root node. The branch node stores the address range of the child node to facilitate navigation to the next layer of child nodes. The leaf node can refer to the bottom node in the tree data structure. The leaf node can refer to the child node of the branch node. The leaf node can store the specific address range and symbol offset of the symbol, wherein the symbol offset refers to the position of the symbol in the symbol table file or the offset relative to the module base address (used to identify the starting position of the module loading position in the memory). Therefore, the target node can refer to the leaf node containing the symbol information of the target address in the tree data structure.
[0105] Reference Figure 3 , Figure 3A structural diagram of a tree data structure in a symbol table file provided for an embodiment of the present disclosure, wherein each branch node of the tree data structure only includes the address of a child node, and the size of each node is fixed, that is, each node manages an address range of a fixed size, and the tree data structure may include a multi-layer structure, which may be divided into a root node, a branch node, and a leaf node, and the size of each node may be fixed at 4KB, and the root node and the branch node only include the starting address (start_addr) of each entry, and each entry occupies 4 bytes, therefore, the root node and the branch node may store 1024 entries, that is, the root node and the branch node may store the starting address of 1024 entries, that is, the branch nodes of the second layer may store a total of 1024*1024 entries, and the order of elements of each entry in the array stored in the node may represent the order of child nodes pointing to the next layer, and the leaf node stores the symbolic offset of the content after address translation, the starting address and the ending address of the corresponding entry, and the symbolic offset (offset), the starting address (start_addr), and the ending address (end_addr) three fields each occupy 4 bytes, 4096 / [s sizeof(offset)+sizeof(start_addr)+sizeof(end_addr)]=341, therefore, each leaf node can store 341 elements, such as Figure 3 As shown in the figure, a tree data structure with a height of 3 can store 341*1024*1024 entry indexes. Since the addresses in the symbol table file are arranged sequentially in the tree data structure, there is no need to compare the target address with the key value of the branch node. The target address can be used to directly locate the branch node that meets the address range of the target address, and then the address is calculated by the address range and element order stored in the branch node to obtain the storage address of the next layer of child nodes in the tree data structure, and locate the corresponding leaf node without relying on the child node pointer for positioning, nor on the linked list structure of the leaf node for range scanning, thereby omitting the data storage space of the child node pointer and the linked list structure. At the same time, compared with the query positioning method using the binary search method, the time complexity of query positioning using the direct calculation method is reduced from O(logn) to O(1), which effectively improves the query efficiency.
[0106] Step 203: Retrieve the function number of the target address and the first line number where the original instruction is located from the first symbol information of the first symbol table file according to the symbol offset.
[0107] Among them, the first symbol table file may include an index file and a data file, wherein the index file is a tree data structure, storing a mapping relationship between an address interval and a symbol offset, and is used to quickly locate the symbol information block where the target address is located, and the data file is the first symbol information, and the first symbol information may refer to the detailed information (function number, line number) of the symbol organized according to the address interval to form multiple symbol information blocks, and can be accessed through the symbol offset, that is, the first symbol information may store the mapping relationship between the symbol offset, the function number and the first line number. The function number may be a sequence number or number used to distinguish and identify the function characters in the running code corresponding to the original instruction, different function numbers indicate different function characters in the original instruction, and the function character may refer to a string representing the function name in the running code of the application, therefore, after obtaining the symbol offset, the function number of the target address and the first line number where the original instruction is located may be retrieved from the first symbol information. Specifically, taking the target address 0x30aa as an example, assume that the corresponding symbol offset is 0x2000 found through the tree data structure of the first symbol table file; according to the symbol offset 0x2000, we can rely on the mapping relationship of the first symbol information to jump to the corresponding symbol information block in the first symbol information, read the content of the symbol information block, and obtain the function number and the first row number of the target address.
[0108] It should be noted that the first line number may refer to the line number of the original instruction in the running code of the application program to identify and locate the original instruction. There may be multiple ways to store the first line number in the first symbol information. For example, a line number array is stored in the symbol information block, and each target address corresponds to a line number. The order of entries in the line number array is consistent with the order of instruction addresses. Assume that the line number array stored in the first symbol information is [10,15]. If the target address is the second instruction in the function, the first line number is 15. For another example, the symbol information block may store the starting line number and offset of the first line number to reduce repeated storage. The starting line number may refer to a certain continuous period. The starting source code line number corresponding to the continuation instruction, the address offset can refer to the offset of the target address relative to the starting address (such as +0x4), and the line number offset can refer to the offset of the line number relative to the starting line number (such as +5), so that a section of instructions with continuous addresses can be mapped to a section of continuous (or jumping) line numbers, and compressed and stored by recording the starting line number and the offset. Therefore, the corresponding line number can be matched by calculating the offset between the target address and the starting address. Assuming the starting line number is 10, when the offset between the target address and the starting address is +0x4, the line number offset corresponding to the address offset +0x4 is +05, so the first line number can be obtained as 15.
[0109] It should be noted that the first line number of the original instruction retrieved from the first symbol information can be a specific value or a range of numbers. Specifically, when the symbol information can clearly identify the specific line code of the original instruction, such as when there is no loop or inline function in the original instruction, the first line number retrieved can be a specific number. In some compilation scenarios, multiple lines of instructions will be merged and reorganized. For example, if a loop or inline function is used in the original instruction, the first line number retrieved at this time is the address range corresponding to these multiple lines of instructions.
[0110] It is worth noting that since the function number is used for storage in the first symbol information, at this time, all addresses of the same function name share a function number, and the function name and row number metadata are only stored once in the first symbol information. The reuse of function numbers can effectively save storage space for function name strings, while reducing repeated storage of different instances of the same function name.
[0111] Step 204: Obtain a first symbol dictionary according to the target address, retrieve the first function character of the original instruction in the first symbol dictionary according to the function number, and obtain the translation result of the original instruction based on the first function character and the first line number.
[0112] Among them, the symbol dictionary can be stored in the symbol table file, and the first symbol dictionary can refer to a data structure used to store the mapping relationship between function characters and function symbols, which is used to map function characters to unique identifiers, namely function numbers, thereby reducing repeated storage and improving query efficiency. Specifically, repeated function names (such as main, prinf, etc.) are stored only once, and subsequently referenced through the corresponding function numbers, which can reduce the size of the symbol table file. At the same time, querying using the function number can locate the corresponding symbol information more quickly, which can effectively reduce overhead compared to querying and locating using the original function name in string form.
[0113] It should be noted that the symbol dictionary can be stored in slices. Specifically, the symbol dictionary can be divided into multiple dictionary slices for storage according to the instruction address. Each address range corresponds to a dictionary slice, that is, the function numbers corresponding to the first function characters used by the original instructions corresponding to an address range are integrated to form a dictionary slice, so that the corresponding dictionary slice can be obtained through the address range where the target address falls, and the function character corresponding to the function number of the target address can be found through the dictionary slice.
[0114] After obtaining the first symbol dictionary, the function number can be used to search and match in the first symbol dictionary to find the first function character corresponding to the function number, and then the obtained first function character is combined with the first line number to obtain the translation result of the original instruction. The translation result of the original instruction refers to the machine code of the original instruction in the running code restored to readable code information. Assuming that the function number of the target address "0x30aa" is "42" retrieved from the above symbol information, and the first line number of the original instruction is "15", the corresponding first symbol dictionary is obtained through the target address "0x30aa", and the first symbol dictionary is searched to read the string "MyClass::crashMethod()" corresponding to the function number "42", that is, the first function character "MyClass::crashMethod()" of the original instruction is obtained, and the first function character "MyClass::crashMethod()" is combined with the first line number "15" to obtain the translation result of the original instruction "MyClass::crashMethod()+15".
[0115] In summary, in the instruction translation method proposed in the embodiment of the present disclosure, the content of the symbol table file is reconstructed, and the first symbol table file contains a tree data structure and the first symbol information. Each node in the tree data structure manages an address range of a fixed size. Therefore, when searching for a node index matching the target address in the tree data structure, the index of the next layer of nodes can be obtained by direct address calculation, without relying on child node pointers for positioning, and without relying on the linked list structure of leaf nodes for range scanning. This not only effectively reduces the size of the symbol table file, but also effectively reduces the time complexity of the query stage. The first symbol information uses function numbers instead of complete function characters for storage, and the tree data structure and the first symbol information are used to complete the establishment of a mapping relationship between the target address, the function number and the first row number. The function number of the target address and the first row number where the original instruction is located are found through the tree data structure, and then the first symbol dictionary obtained by the target address is used to retrieve the first function character corresponding to the function number, and the first function character is combined with the first row number to restore the translation result of the original instruction, effectively improving the efficiency of instruction translation.
[0116] In a possible implementation, the instruction translation request may include information related to the target address of the original instruction, such as an offset address, which refers to the machine instruction address pointed to by the program counter when the application crashes. At the same time, the instruction translation request also includes module information, such as the name of the module where the crash occurs, the unique identifier of the module, and the loading base address of the module in the memory, wherein the target address of the original instruction may refer to the actual address offset, that is, the address information obtained by subtracting the loading base address from the offset address, wherein the translation process of the original instruction depends on the corresponding symbol table file, which is usually uploaded to the server for storage. The instruction translation request usually includes the following field data: service object identifier, event occurrence time, event type (such as crash event, freeze event, etc.), platform type (such as xx operating system), application version number, message unique identifier and stack information, and the purpose of the instruction translation request is to translate the stack information using the symbol table file and restore it into readable code information. The service object identifier is the unique identifier of the service object. The symbol table file of each service object is independently stored in the server through the service object identifier to ensure data isolation and security. The service object can be a specific program product. The event occurrence time and event type are used to indicate that the application triggers the client SDK to generate instructions. The event type and time of the command translation request. The event type can be used to identify the type of event that occurs in the current application. For example, the event type "crash" indicates that the application has crashed. The platform type and application version number are used to distinguish between operating systems or hardware architectures, as well as different versions of the same application. The instruction sets and debugging formats of different platforms are different, and independent symbol table files are required. After the application has undergone version iterations and the application code has been updated, the corresponding version of the symbol table file is also required. Therefore, these meta-information can assist in determining the symbol table file required for the command translation request. In addition, each generated command translation request will have a corresponding message unique identifier, which is convenient for the server to detect the status of the command translation request.It is worth noting that the stack information may include module name, module base address, offset address, symbol table file identifier, etc. For example, the module name is "AAA" to indicate the code module (such as a dynamic link library, executable file) where the jam occurs; the module base address is "0x7fff60302000", indicating the starting address of the current module "AAA" loaded into the memory; the offset address is "0x7fff60304950", indicating the offset of the running address where the jam occurs relative to the base address, and the symbol table file identifier is "a4938cf5". Therefore, the corresponding symbol table file can be directly searched from the server through the symbol table file identifier "a4938cf5"; and since the symbol table file is stored independently according to each service object identifier, the storage path of the symbol table file can be constructed using the service object identifier, platform type, application version number and symbol table file identifier, and the corresponding symbol table file can be located through the storage path, and then the symbol table file can be used to translate the original instruction of the stack information.
[0117] In one possible implementation, referring to Figure 4 , Figure 4A flow chart of the instruction translation service provided for the embodiment of the present disclosure, the client SDK can capture meta-information such as module name, base address, offset address, etc. when a crash occurs, and then automatically associate the corresponding symbol table file identifier according to the module name, application version information, platform type, etc., integrate the module name, base address, offset address and symbol table file identifier, etc. into stack information, and then generate a unique message identifier for this crash, and then integrate data such as event type, stack information, platform type, application version number, etc., and encapsulate them into an instruction translation request. Sensitive information can be encrypted if necessary, and then the instruction translation request is sent to the traffic gateway via the https protocol. The traffic gateway will distribute the instruction translation request to the available symbol translation service node, i.e., the server. The traffic gateway will verify the instruction translation request, such as checking the field integrity and verifying the legality of the request format, and the traffic gateway can filter duplicate instruction translation requests. After the server receives the instruction translation request, it will parse the instruction translation request, first extract the module base address and offset address from the stack information, calculate the target address, and then retrieve the translation result of the original instruction corresponding to the target address in the local cache. If the translation result of the original instruction corresponding to the target address exists in the local cache, the translation result is directly returned. If the translation result of the original instruction corresponding to the target address does not exist in the local cache, the symbol table file identifier parsed from the stack information can be used to download the first symbol table file from the cloud object storage to the local disk, or the service object identifier, platform type, application version number and symbol table file identifier parsed from the instruction translation request are used to construct a storage path for the symbol table file, and the first symbol table file is downloaded from the cloud object storage according to the storage path; after the first symbol table file is saved to the local disk, the first symbol table file can be used to translate the original instruction of the stack information to obtain the translation result of the original instruction corresponding to the target address. Among them, the first symbol table file is obtained by reconstructing the original symbol table file. After obtaining the first symbol table file, the symbol table meta-information is extracted from the reconstructed first symbol table file, and then the first symbol table file is uploaded to the cloud object storage for storage according to the symbol table meta-information.
[0118] In one possible implementation, when uploading the original symbol table file of the application to the server, a content reconstruction process can be introduced. The original symbol table file may refer to a data structure file generated during the compilation or linking process without content reconstruction. This is because the original symbol table file corresponding to the application stores other information unrelated to instruction translation in addition to debugging information, and uses a complex data structure to describe the mapping relationship, resulting in the original symbol table file being too large. By introducing the content reconstruction process, irrelevant information in the original symbol table file is trimmed to reduce the file size. At the same time, the indexing process is optimized, and a symbol dictionary is additionally introduced for instruction translation. The optimized instruction translation method can adapt to symbol table files of different platforms and different file formats, thereby improving the applicability of the instruction translation method.
[0119] Specifically, please refer to Figure 5 , Figure 5 The schematic diagram of the structure of the symbol table file and the symbol dictionary proposed in the embodiment of the present disclosure is as follows. The symbol table file mainly includes the following three parts: header information, tree data structure (index file) and symbol information (data file). The header information is used to record the meta information of the symbol table file. Specifically, the header information can store predefined identifiers to identify the data formats of different symbol table files, such as DWARF format and PDB format. It can also distinguish between so files and dYSM files in the DWARF format. At the same time, the header information also includes the version information of the symbol table file (such as Figure 5 The main version information and sub-version information shown in the figure), as well as the index size (volume size), number of layers, index starting offset, layer offset and number of elements of each layer structure, leaf starting offset of leaf nodes, symbol starting offset of symbol information, symbol dictionary starting offset of symbol dictionary, etc. of the tree data structure, so as to facilitate subsequent updates to the symbol table file.
[0120] Among them, the tree data structure is used to record the index of symbol information, and the corresponding symbol information is retrieved through the index of the target address. The tree data structure divides nodes according to the address range, and each node records the starting address of the corresponding address range. In the tree data structure, the size of all nodes is fixed, and the non-leaf nodes of the tree data structure, namely the root node and the branch node, only save the 4-byte starting address (start_addr of uint32_t type). The starting addresses stored in the root node and the branch node can be arranged in ascending order. Therefore, both the root node and the branch node can store 1024 entries, namely 1024 starting addresses. The leaf node stores a 4-byte starting address (uint32_t type start_addr), an ending address (uint32_t type end_addr) and a symbol offset (uint32_t type offset), that is, the leaf node stores 341 elements; the root node is the first layer of the tree data structure, so the address of the root node in the symbol table file is the index starting offset + the first layer offset, and the second layer of the tree data structure includes multiple branch nodes, and the address of the first branch node in the second layer in the symbol table file is the index starting offset + the second layer offset Shift, and the address of the second branch node in the second layer needs to be the address of the first branch node in the second layer plus a 4KB address offset, the third node in the second layer needs to be the address of the first branch node in the second layer plus two 4KB address offsets, and so on; similarly, the leaf starting offset corresponds to the address of the first leaf node in the symbol table file, and adding a 4KB address offset on the basis of the leaf starting offset will obtain the address of the second leaf node in the symbol table file, and so on, so the root node and the branch node do not need to store child node pointers, and can locate the child nodes of the next layer by directly calculating the address offset. This search method not only optimizes the pointer storage space, maximizes the number of child nodes of each node, so as to reduce the height of the tree data structure and compress the size of the symbol table file, but also, compared with the binary search method to search for a symbol offset corresponding to an address in a symbol table file, logN disk read operations are required, and the search method adopted in the instruction translation method proposed in the embodiment of the present disclosure does not need to perform key value comparison, and only 5 disk read operations are required to complete a stack search, reducing the time complexity of the indexing process from O(logN) to O(1).
[0121] Among them, the symbol information is a data file, which is used to record the function number, line number, and whether it is an inline function and other information corresponding to each target address. The symbol information can be divided according to the address range of the target address to form multiple symbol information blocks. The symbol information and the tree data structure are used to construct a mapping relationship between the target address, function number and line number. The symbol offset provided by the target node in the tree data structure can be used to locate the corresponding symbol information block, so that the function number of the target address and the first line number of the original instruction can be extracted from the symbol information block.
[0122] Among them, the symbol dictionary can store a data structure of the mapping relationship between function numbers and function characters. After obtaining the function number through symbol information, the corresponding first symbol dictionary can be obtained through the target address. Using the mapping relationship between function numbers and function characters stored in the first symbol dictionary, the function character corresponding to the function number to be translated is retrieved and the function character is used as the first function character of the original instruction.
[0123] It is worth noting that the symbol information stores the function number, which can save storage space compared to the way the original symbol table file stores function characters. Specifically, assume that the original symbol table file stores the following symbol information code:
[0124] AbcDefg::HijkLmn()(OpqRst..uv:249)
[0125] AbcDefg::HijkLmn()(OpqRst..uv:250)
[0126] AbcDefg::HijkLmn()(OpqRst..uv:251)
[0127] It can be seen that the symbol information in the original symbol table file records the complete class characters, function characters and line numbers, and in the instruction translation method proposed in the embodiment of the present disclosure, by introducing a symbol dictionary, the class character "AbcDefg::HijkLmn()" is represented by the class number "1", and the function character "OpqRst..uv" is represented by the function number "2". Therefore, the symbol information codes recorded in the first symbol table file correspond to "1 2 249", "1 2 250", and "1 2 251", thereby greatly reducing the repeated function characters and class characters. It should be noted that, in addition to function characters, some original instructions may also contain class characters. In the symbol table file, class characters may adopt the same optimization strategy as function characters. The class characters are represented by unique class numbers, and the class numbers are used instead of the class characters to be stored in the symbol table file. Specifically, the class numbers may share the same storage structure with the function numbers. For example, the symbol dictionary also stores the mapping relationship between the class numbers and the class characters. Thus, by assigning unique class numbers to the class characters and sharing the same retrieval framework with the function numbers, the redundant storage overhead can be significantly reduced while maintaining the efficient retrieval capability of the symbol table file.
[0128] In summary, by reconstructing the content of the symbol table file, introducing a symbol dictionary, and adopting a hierarchical index design of a tree data structure in the symbol table file, the efficiency of symbol information storage and retrieval is optimized. Specifically, by assigning corresponding class numbers and character numbers to class characters and function characters for storage, the occurrence of repeated character strings is reduced, and the size of the symbol table file is effectively reduced. At the same time, the tree data structure uses hierarchical nodes of fixed size, which can replace pointer storage with address calculation, maximize the use of disk block space, further compress the size of the symbol table file, and reduce the time complexity of the retrieval process, achieving fast response in high concurrency scenarios. In addition, the header information also uses a custom identifier to identify symbol table files of different formats, which, combined with the platform-independent index structure, can be compatible in a multi-platform environment.
[0129] In a possible implementation, the first symbol table file can be obtained by reconstructing the original symbol table file. Specifically, the original symbol table file can be first obtained, and multiple target addresses, first function characters corresponding to each target address, and first row numbers corresponding to each first function character can be extracted from the original symbol table file; then, a tree data structure is created based on the multiple target addresses and symbol offsets, a function number is assigned to the first function character, and first symbol information is created based on the symbol offset, the function number, and the first row number, and the tree data structure and the first symbol information are encapsulated as a first symbol table file; then, a first symbol dictionary is created based on the function number and the first function character, thereby completing the reconstruction of the first symbol table file, filtering out irrelevant information in the original symbol table file, and reducing the size of the symbol table file.
[0130] In a possible implementation, the original symbol table file may refer to the file generated when the application is compiled and built, and the original symbol table file may be a dSYM file of an iOS / macOS platform, or an so file of an Android platform, or a pdb file of a Windows platform, or an so file of a Linux platform. By parsing the original symbol table file, all debugging information such as target addresses (instruction addresses), function characters, class characters, and line numbers are extracted, and then irrelevant metadata such as unused symbols, debug format descriptors, etc. are filtered, wherein the extracted data may be grouped according to the target address, which is equivalent to combining the target address, the first function character corresponding to the target address, and the first line number corresponding to the first function character into a group of address data pairs, specifically, a group of address data pairs may be expressed as (addr1, func1, line1), that is, the target address addr1, the first function character func1 and the first line number line1 correspond respectively, and multiple groups of address data pairs may be extracted from the original symbol table file, so as to facilitate the subsequent filling of data when reconstructing the symbol table file.
[0131] In a possible implementation, in the process of constructing a tree data structure, all target addresses can be first arranged in ascending order and divided into multiple continuous intervals. For example, each 4KB address space is used as an address interval, that is, each node is fixed in size of 4KB. Then, the starting addresses of all address intervals (such as 0x1000, 0x2000, 0x3000, etc.) are stored in the address space of the root node. Each starting address, that is, entry occupies 4 bytes. The root node can have a total of 1024 entries. Then, each address interval is subdivided, and the starting addresses of the subdivided address sub-intervals are stored in the address space of the branch node. The structure of the branch node is the same as that of the root node. Among them, multiple branch nodes can be divided into multiple layers of sorting, and the address sub-intervals of the upper branch nodes are sorted. The interval is further subdivided, and the starting address of the upper-level address interval after subdivision is stored in the address space of the lower-level branch node, and so on; after completing the construction of the branch nodes of the tree data structure, the leaf nodes are constructed. In addition to storing the starting addresses of each address interval, the address space of the leaf nodes also stores the ending address of the address interval and the symbol offset corresponding to the address interval. The starting address, the ending address and the symbol offset each occupy 4 bytes, so each entry occupies 12 bytes, which is equivalent to a leaf node that can store 341 entries. It should be noted that the symbol offset stored in the leaf node has an indexable association with the first symbol information. The symbol offset refers to the offset address of the symbol information block corresponding to the target address stored in the first symbol information.
[0132] In one possible implementation, a globally unique function number is assigned to each unique function character. Similarly, a globally unique class number can also be assigned to each unique class character. A symbol dictionary is then generated according to the mapping relationship between function characters and function numbers, and the mapping relationship between class characters and class numbers. Specifically, the symbol dictionary can store these two mapping relationships in the form of a hash table or an array.
[0133] In one possible implementation, the function characters in the address data pair are replaced with the function number, and then the replaced address data pair, i.e., the target address, function number, and row number, are packaged into a symbol information block of a fixed size according to the address range. At the same time, the address range, i.e., the symbol offset corresponding to the target address, is associated with the corresponding symbol information block, so that the symbol information block carries the symbol offset of the tree data structure, and an index relationship between the tree data structure and the symbol information is constructed. Then, continuous target addresses with the same function number can be merged, such as multiple target addresses of an inline function share the same function number, thereby optimizing the storage of the symbol information block. Then, all the symbol information blocks are combined to construct the symbol information.
[0134] In one possible implementation, after completing the construction of the tree data structure and the symbol information, the tree data structure and the symbol information can be encapsulated to obtain a reconstructed symbol table file, wherein a predefined identifier can be used to identify the format of the original symbol table file, the structural information of the tree data structure, and the index information of the symbol information to form the header information of the symbol table file. In addition, the symbol dictionary corresponding to the symbol information can also be encapsulated in the symbol table file.
[0135] In one possible implementation, the user uploads the original symbol table file to the server, and the server reconstructs the original symbol table file, recreates the tree data structure and the first symbol information based on the extracted target address, the first function character, the first line number and other symbol table meta-information, and then encapsulates the tree data structure and the first symbol information to obtain a new first symbol table file, and additionally creates a first symbol dictionary, and associates the first symbol dictionary with the first symbol table file, so that the first symbol dictionary and the first symbol table file can work together in the subsequent instruction translation process.
[0136] In one possible implementation, the process of using the target address to retrieve the target node in the tree data structure can specifically be to first retrieve a first reference address that matches the target address in the root node of the tree data structure of the first symbol table file; then, retrieve a second reference address that matches the target address in the branch node associated with the first reference address; and then, determine the leaf node associated with the second reference address as the target node.
[0137] The root node stores multiple start address entries, which are arranged in ascending order. The address range corresponding to each entry can be obtained by the start address of the next adjacent entry. Specifically, the start address of the current entry is used as the start address of the address range, and the start address of the next adjacent entry is used as the end address of the address range. Figure 6 , Figure 6 A schematic diagram of the indexing process of the tree data structure provided for an embodiment of the present disclosure, assuming that the starting address of the third entry in the root node is 0x103000, and the starting address of the fourth entry is 0x104000, then the address range represented by the third entry is [0x103000, 0x104000), that is, the root node stores multiple address ranges, namely, multiple continuous address intervals formed by all target addresses being arranged and divided in ascending order.
[0138] Among them, the process of searching the target address in the root node is specifically to determine which starting address entry's address range the target address falls into, and the first reference address may refer to the starting address corresponding to the specific address range that the target address falls into in the root node, and the second reference address may refer to the starting address corresponding to the address range that the target address falls into in the branch node. It should be noted that the address range calculation method of each entry in the branch node is the same as the address range calculation method of each entry in the root node, and the process of retrieving the second reference address that matches the target address in the branch node is the same as the process of retrieving the first reference address that matches the target address in the root node.
[0139] Among them, when there are multiple layers of branch nodes in the tree data structure, after retrieving the first reference address matching the target address in the root node, the intermediate reference address matching the target address is retrieved in the branch node associated with the first reference address. The intermediate reference address also refers to the starting address corresponding to the address range within which the target address falls in the branch node. All address ranges of the next layer of branch nodes are obtained based on the subdivision of one of the address ranges of the previous layer of branch nodes. Therefore, the intermediate reference address can be associated with the next layer of branch nodes; when the current layer of branch nodes is the last layer of branch nodes, the second reference address matching the target address can be retrieved in the branch node associated with the intermediate reference address, and then the leaf node associated with the second reference address is determined as the target node.
[0140] like Figure 6As shown, assuming that the target address is 0x103388, which falls within the address range of the third entry [0x103000, 0x104000), the starting address 0x103000 of the address range of the third entry is the first reference address that matches the target address. After determining the first reference address that matches the target address in the root node, find the branch node associated with the first reference address according to the index method of the tree data structure. Assume that the first reference address 0x103000 finds the associated branch node C, where the structure of the branch node and the root node are the same, and the branch node also stores multiple starting address entries, but the starting address stored in the branch node is the starting address of the sub-range obtained by subdividing the address range of the upper layer node, which is equivalent to subdividing the address range of the third entry of the root node [0x103000, 0x104000) into multiple sub-ranges, and the starting addresses of these sub-ranges can be stored in the address space of the branch node C. It should be noted that the address range of the entry in the upper node can be divided into multiple sub-ranges, and the starting addresses of the multiple sub-ranges are stored in the lower sub-nodes in ascending order of address, and these starting addresses can be stored in the same lower sub-node. Then, the second reference address matching the target address can be retrieved in the branch node. Assuming that the target address 0x103388 falls into the address range [0x103300, 0x103500) of the second entry in branch node C, then the starting address 0x103300 of the address range of the second entry in branch node C is the second reference address matching the target address. Then, the leaf node associated with the second reference address is found according to the indexing method of the tree data structure, and the associated leaf node is used as the target node. Therefore, the initial address of each node can be used for positioning, without comparing the key value of the target address and the branch node. The root node and the branch node do not need to store the child node pointer, which can effectively reduce the time complexity of the index and optimize the volume of the tree data structure.
[0141] In one possible implementation, the indexing method of the tree data structure may refer to determining the element position of the reference address associated with the target address in the current node, determining the node address offset based on the element position, then obtaining the node starting address of the leaf node associated with the current node, determining the destination node address based on the node starting address and the node address offset, and then finding the corresponding leaf node based on the destination node address. Therefore, by combining the node size and the pre-stored node starting address, the physical position of any node can be directly calculated, eliminating the pointer storage space.
[0142] Among them, if the current node is the root node, the reference address associated with the target address is the first reference address, the element position may refer to the order of entries of the first reference address in the root node, the node address offset refers to the relative offset of the first branch node in the second layer of the branch node corresponding to the first reference address, the node start address may refer to the address of the layer offset of the second layer of the tree data structure plus the index start offset, which is equivalent to the starting address of the tree data structure in the symbol table file plus the starting address of the first node in the second layer of the tree data structure, and the destination node address refers to the specific address of the leaf node corresponding to the reference address in the symbol table file. Specifically, the node address offset can be calculated by multiplying the element position by the node size, such as Figure 6 As shown, assuming that the node size is 4KB, it is necessary to locate the branch node associated with the first reference address 0x103000. The first reference address is the starting address of the third entry in the root node, that is, the element position of the first reference address in the root node is 3. The value after subtracting one from the element position is multiplied by the node size to obtain a node address offset of 0x2000. Then, the node starting address is added to the node address offset to obtain the destination node address.
[0143] Among them, if the current node is a branch node, the reference address associated with the target address is the second reference address, the element position refers to the order of entries of the second reference address in the branch node, the node address offset refers to the relative offset of the leaf node corresponding to the second reference address to the first leaf node in the same layer, and the node starting address refers to the starting address of the first node in the layer where the leaf node is located in the tree data structure, that is, the address obtained by adding the leaf starting offset in the header information to the index starting offset; the node address offset at this time needs to be calculated in combination with the element position, the node position of the current branch node in the layer, and the number of entries stored in the branch node. Specifically, the value after the node position is subtracted by one is multiplied by the number of entries, and then the value after the element position is subtracted by one is added to the product result, and then the sum is multiplied by the node size to obtain the node address offset, and then the node starting address is added to the node address offset to obtain the destination node address.
[0144] Among them, if the tree data structure includes multiple layers of branch nodes, when addressing the address of the next layer of branch nodes in the upper layer branch node, determine the element position of the intermediate reference address in the current branch node, and then determine the node position of the current branch node in the layer where it is located, and determine the node address offset according to the element position, node position and number of entries. Similarly, multiply the value obtained by subtracting one from the node position by the number of entries, and then add the value obtained by subtracting one from the element position to the product structure, and then multiply the sum by the node size to obtain the node address offset, and then add the node starting address to the node address offset to obtain the destination node address. It should be noted that the node starting address at this time is obtained by adding the index starting offset to the layer offset corresponding to the next layer of branch nodes. The index starting offset and the layer offset of the corresponding layer can be read through the header information of the symbol table file.
[0145] In one possible implementation, the node size of the tree data structure can be aligned with the memory block size of the operating system to ensure that each disk read operation can efficiently obtain complete node data. The memory block size can refer to the memory page size (Page Size), which can refer to the basic unit of memory management in the operating system. Since different operating systems and hardware architectures may use different memory block size configurations, the system identifier of the operating system from which the instruction translation request comes can be obtained from the instruction translation request, and the memory block size corresponding to the operating system can be determined based on the system identifier; then, the node address offset is determined based on the product of the element position and the memory block size.
[0146] Specifically, the system identifier can be parsed and extracted from the instruction translation request. The system identifier can refer to information used to uniquely identify the type of operating system. The system identifier can refer to the platform type in the instruction translation request of the above embodiment. For example, the memory block size of the Linux operating system is 4KB, and the memory block size of the Windows operating system is 4KB. When it is recognized that the system identifier indicates the Linux operating system or the Windows operating system, the node size of the tree data structure can be set to 4KB accordingly, and the memory block size of some special operating systems is 16KB. If it is recognized that the system identifier indicates this part of the special operating system, the node size of the tree data structure can be set to 16KB accordingly, so that a single disk read can obtain complete node data, reduce fragmented access, avoid additional disk read operations caused by cross-page reading, and dynamically adapt the memory block size through the system identifier to support the difference in memory block size of different operating systems.
[0147] It is worth noting that if the system identifier does not match, that is, the system identifier indicates an unknown operating system, a memory block size of 4KB may be used by default and a warning log may be recorded.
[0148] It should be noted that when constructing a tree data structure, it is necessary to identify the operating system corresponding to the original symbol table file, determine the memory block size of the operating system to be applied, and then adjust the node size of the tree data structure according to the memory block size so that the node size of the tree data structure can be aligned with the memory block size of the operating system.
[0149] It is worth noting that the node address offset is determined according to the product of the element position and the memory block size. Specifically, the value of the element position minus one is multiplied by the memory block size to obtain the address offset. For example, assuming that the system identifier indicates the Linux operating system, it is determined that the memory block size is 4KB, and the node address offset corresponding to the third element is (3-1)×4KB, that is, 8KB. At this time, the physical address of the node associated with the third element is the sum of the index start offset, the layer offset of the corresponding layer, and 8KB; for another example, assuming that the system identifier indicates the AIX operating system, it is determined that the memory block size is 64KB. At this time, the node address offset corresponding to the third element is (3-1)×64KB, that is, 128KB. In the AIX operating system, the physical address of the node associated with the third element is the sum of the index start offset, the layer offset of the corresponding layer, and 128KB.
[0150] In a possible implementation, when translating the stack information in the instruction translation request, the order of magnitude of the target address that needs to be translated is much smaller than the order of magnitude of the target address recorded in the full symbol table file. Specifically, in the crash log analysis scenario, dozens of target addresses usually need to be translated, while the full symbol table file may store hundreds of millions of target addresses. If the symbol table file is downloaded in full, it will waste bandwidth and affect the query efficiency. Therefore, a symbol table file sharding strategy can be adopted to split the symbol table file into multiple file shards according to the address range, i.e., the first symbol table file, for example, each shard manages an address interval of 0x10000, wherein the volume of the file shards (i.e., the first symbol table file) split from the full symbol table file can be adapted to the memory block size of the specific operating system, and at the same time, indexes are constructed for these split first symbol table files, and information such as the address range of the file shards is independently stored to facilitate rapid positioning of the target file shards. For example, the constructed index can be the first symbol table file identifier, so that when translating instructions, the corresponding symbol table file shards are directly obtained according to the first symbol table file identifier, reducing the amount of data transmission. It is worth noting that when the symbol table file is stored in fragments, the tree data structure and the symbol information need to be split according to the address range. The first symbol table file identifier records the address range after fragmentation, the storage address of the symbol information after fragmentation, and the storage address of the tree data structure after fragmentation, that is, the mapping relationship between the storage target address, the tree data structure and the symbol information.
[0151] Assuming that the total address range of the target address recorded in the full symbol table file is 0x00000000 to 0xFFFFFFFF, with a total volume of 4GB, the full symbol table file can be split into 100 symbol table file fragments, namely the first symbol table files, each fragment manages the address space of 0x04000000, and the size of each file fragment is 64MB. These first symbol table files are uploaded to the cloud object storage, and the first symbol table file identifier of each first symbol table file can indicate that the corresponding address range is stored. Therefore, when the target address 0x10 3388, the first symbol table file identifier indicating the address range of 0x100000 to 0x140000 can be found, and the first symbol table file identifier can be written into the instruction translation request, so that the server or terminal can extract the first symbol table file identifier from the instruction translation request during the instruction translation process, and locate and download the corresponding first symbol table file in the cloud object storage through the first symbol table file identifier, without downloading the full symbol table file, thereby realizing on-demand loading of symbol table files and reducing invalid data cache. At the same time, the sharding strategy can effectively reduce translation delay.
[0152] In a possible implementation, when the first symbol table file is stored in fragments, the symbol dictionary can also be split. Specifically, the service object identifier of each service object, the first original dictionary of each service object, and the first symbol table file identifier can be obtained; then the first original dictionary is divided into multiple first candidate dictionaries, and a corresponding symbol dictionary index is assigned to each first candidate dictionary; then, the mark file, the first symbol table file identifier, the service object identifier, multiple symbol dictionary indexes, and multiple first candidate dictionaries are stored in an associated manner.
[0153] In a possible implementation, the first original dictionary may refer to a data structure that records the mapping relationship between all function numbers and all function characters that appear in the first symbol table file. The first original dictionary may be split into multiple first candidate dictionaries according to the function number. The first candidate dictionary may refer to a fragment file of the first original dictionary, storing the mapping relationship between some function numbers and function characters, and each first candidate dictionary may correspond to a symbol table file, and the mapping relationship between the symbol table file and the first candidate dictionary may be recorded in the mark file, indicating that the function number in the symbol information of the symbol table file can be mapped and translated into the function character through the corresponding first candidate dictionary. Specifically, the first original dictionary can be divided according to the function number of the first symbol table file corresponding to the first symbol table file identifier, and the first symbol table file can be sequentially divided into the first candidate dictionary. The first function number and the corresponding first function character appearing in the file are written as entries into the next first candidate dictionary until the storage space occupied by the first candidate dictionary reaches a preset threshold, and the next first candidate dictionary is created to continue writing entries until all the first function numbers and the corresponding first function characters appearing in the first symbol information segment are written, wherein the preset threshold may refer to the upper limit of the dictionary file storage volume, which is used to limit the volume of the candidate dictionary, and then a corresponding symbol dictionary index is assigned to each first candidate dictionary. The symbol dictionary index may indicate an address range, which is composed of the target addresses corresponding to all the function numbers recorded in the candidate dictionary, and then the tag file, these symbol dictionary indexes, the corresponding first candidate dictionary, the first symbol table file identifier, and the service object identifier are stored in association. Reference Figure 7 , Figure 7 A schematic diagram of the effect of associating the symbol dictionary and the symbol table file provided in the embodiment of the present disclosure. The service object can upload multiple symbol table files to the cloud object storage for storage. The cloud object storage will classify the files according to the service object identifier and further classify the files under the same service object identifier according to the first symbol table file identifier, such as Figure 7As shown, the first service object stores two first symbol table files in the cloud object storage, and the corresponding first symbol table file identifiers are symbol table 1 and symbol table 2, respectively, wherein different first symbol table files can be associated with different first original dictionaries, for example, the first original dictionary 1 is split according to the function number appearing in the first symbol table file of symbol table 1, and the splitting obtains the first candidate dictionary 1, the first candidate dictionary 2, and the first candidate dictionary 3. Therefore, these first candidate dictionaries, the corresponding symbol dictionary index 1, symbol dictionary index 2, symbol dictionary index 3, and the mark file can be stored in a directory created with the first symbol table file identifier symbol table 1. Similarly, with the service object identifier as the first-level directory and the symbol table 2, the first symbol table file identifier, as the second-level directory, the first original dictionary 2 associated with the first symbol table file of the symbol table 2 is split into the first candidate dictionary 4, the first candidate dictionary 5, and the first candidate dictionary 6. Therefore, the first candidate dictionary 4, the first candidate dictionary 5, the first candidate dictionary 6 and the corresponding symbol dictionary index and mark file association are stored in the second-level directory of the symbol table 2. Therefore, by splitting the first original dictionary into multiple first candidate dictionaries, the symbol dictionary is stored in fragments, which helps to flexibly store data and optimize resource utilization.
[0154] In one possible implementation, the first original dictionary can be split according to the first symbol table file segment, indicating that the function number appearing in the symbol information of the first symbol table file segment can be mapped and translated by the corresponding first candidate dictionary. Similarly, the first function number and the corresponding first function character appearing in the first symbol table file segment are sequentially written as entries into the next first candidate dictionary until the storage space occupied by the first candidate dictionary reaches a preset threshold, and the next first candidate dictionary is created to continue writing entries until all the first function numbers and the corresponding first function characters appearing in the first symbol table file segment are written. Then, a corresponding symbol dictionary index is assigned to each first candidate dictionary, and the mark file, these symbol dictionary indexes, the corresponding first candidate dictionary, the first symbol table file segment, and the service object identifier are stored in association.
[0155] In one possible implementation, after obtaining the function number, when it is necessary to use a symbol dictionary to translate the function characters, you can first obtain the tag file, retrieve the symbol dictionary index in the tag file according to the target address, and then obtain the first symbol dictionary from multiple first candidate dictionaries stored in the shards according to the first symbol table file identifier and the symbol dictionary index.
[0156] Among them, refer to Figure 8 , Figure 8A process diagram of symbol dictionary indexing provided for an embodiment of the present disclosure, wherein a tag file may refer to a data structure recording a mapping relationship between multiple address ranges and multiple symbol dictionary indexes, to indicate that the first candidate dictionary can translate function characters for function numbers within the address range, which is equivalent to the symbol dictionary index of each first candidate dictionary indicating an address range, which is composed of target addresses corresponding to all function numbers recorded in the candidate dictionary, wherein the address range may be an explicitly expressed address range, that is, the address range contains a start address and an end address, such as Figure 8 As shown, the tag file records three address ranges and corresponding symbol dictionary indexes, namely address range 1, address range 2 and address range 3. Address range 1 is represented as "0x1:0x5000", with a starting address of "0x1" and an ending address of "0x5000"; address range 2 is represented as "0x5001:0x10000", and address range 3 is represented as "0x10001:0x15000". Address range 1 corresponds to symbol dictionary index A, address range 2 corresponds to symbol dictionary index B, and address range 3 corresponds to symbol dictionary index C. Therefore, the target address can be used to retrieve the address range where the target address is located in the tag file, and the corresponding symbol dictionary index can be determined, that is, the symbol dictionary index can represent the address range where the target address is located, and then the first symbol table file identifier can be used to find multiple first candidate dictionaries stored in the fragments, and then the symbol dictionary index can be used to locate the first symbol dictionary, such as Figure 8 As shown, the first symbol dictionary 1 can be located through the symbol dictionary index A, the first symbol dictionary 3 can be located through the symbol dictionary index B, and the first symbol dictionary 2 can be located through the symbol dictionary index C. Assuming that the target address is "0x10338", the symbol dictionary index C indicates the address range where the target address "0x10338" is located. The symbol dictionary index C is used as the symbol dictionary index required for the target address "0x10338", and the first symbol dictionary 3 is located according to the symbol dictionary index C. Then, the function number corresponding to the target address "0x10338" is translated using the first symbol dictionary 3. Therefore, the required dictionary slice can be directly located through the symbol dictionary index without loading the full symbol dictionary, thereby realizing on-demand loading of slices and reducing invalid data transmission.
[0157] It should be noted that the address range indicated by the tag file and the symbol dictionary index can be an implicitly expressed address range, that is, it only indicates a starting address or an ending address. The address range is calculated for the starting address or the ending address through a preset address calculation method to obtain the address range corresponding to the symbol dictionary index, and then it is determined which address range indicated by the symbol dictionary index the target address falls into.
[0158] In one possible implementation, a single address range recorded in the tag file may exceed the address range indicated by a symbol dictionary index. In this case, when translating the function number of the single address range recorded in the tag file, multiple first symbol dictionaries are required, that is, the single address range recorded in the tag file can indicate multiple symbol dictionary indexes at the same time.
[0159] In a possible implementation, the tag file may store multiple address ranges and file index paths, where the file index path may refer to the storage address of the tree data structure on the cloud object storage, the storage address of the symbol information on the cloud object storage, and the symbol dictionary index. Fig. 9 , Fig. 9 A schematic diagram of the process of symbol dictionary indexing provided for another embodiment of the present disclosure. Similarly, the tag file record has three address ranges and corresponding file index paths, namely address range 1, address range 2 and address range 3. Address range 1 is represented as "0x1:0x5000", address range 2 is represented as "0x5001:0x10000", and address range 3 is represented as "0x10001:0x15000". After receiving the instruction translation request, the target address is extracted from the instruction translation request, and the address range in which the target address falls in the tag file is determined. The corresponding tree data structure, symbol information and symbol dictionary index are located according to the file index path corresponding to the address range, and then the function number of the target address is retrieved using the tree data structure and symbol information. The corresponding first symbol dictionary is then found to translate the function number of the target address. Assuming that the target address extracted from the instruction translation request is "0x888" and falls within address range 1, the tree data structure, symbol information and symbol dictionary index are located according to the file index path corresponding to address range 1, and the symbol offset of the target address "0x888" is retrieved according to the tree data structure and symbol information, where the file index path may contain multiple symbol dictionary indexes, and the function number corresponding to address range 1 requires multiple first symbol dictionaries for translation, such as Fig. 9 As shown, the file index path corresponding to the address range 1 contains symbol dictionary index 1 and symbol dictionary index 2. Therefore, the first symbol dictionary 1 and the first symbol dictionary 2 are located and downloaded according to symbol dictionary index 1 and symbol dictionary 2, and then the first symbol dictionary 1 and the first symbol dictionary 2 are used to translate the function number of the target address "0x888".
[0160] In a possible implementation, in order to achieve a balance between storage efficiency, retrieval performance and resource overhead, especially when processing massive symbol data, such as hundreds of millions of symbol table files, a multi-level storage method can be used to store the first symbol table file. Because the symbol table file in some large projects (such as operating system kernels, game engines) is too large, loading the full amount into memory or single-layer storage will take up too many resources, and in actual debugging or crash analysis, the number of target addresses that need to be translated is relatively small (such as dozens or hundreds), but the full symbol table file covers all possible target addresses, and loading the full symbol table file will cause a waste of resources. Therefore, the full symbol table file is fragmented to form a first symbol table file and then stored hierarchically. Specifically, the first symbol table file can be stored in a local disk, or a network file system, or a cloud object storage, or an online analytical processing database (Online Analytical Processing, OLAP), wherein a cold and hot data hierarchical cache strategy can be used to store the current hot spot first symbol table file in the local disk, and the cold data can be removed from the local disk and saved in the network file system or cloud object storage.
[0161] In a possible implementation, when the first symbol table file needs to be called, the first symbol table file can be first obtained from the local disk according to the first symbol table file identifier. When the first symbol table file does not exist in the local disk, the first symbol table file is obtained from the network file system according to the first symbol table file identifier; when the first symbol table file does not exist in the network file system, the first symbol table file is obtained from the cloud object storage according to the first symbol table file identifier. Such a multi-level storage calling process is because in the multi-level storage structure of the first symbol table file, the data is sorted according to the access speed and cost. The files stored in the local disk can achieve fast access but high cost. The network file system can achieve shared storage in the local area network with medium access speed, and the cloud object storage can achieve low-cost archival storage, but the access delay is the highest. Therefore, the file is searched from the high-speed storage first, and if it is not hit, it will fall back to the low-speed storage in turn. Through the multi-level fallback mechanism, the calling process of the symbol table file can reduce the storage cost while ensuring performance.
[0162] In one possible implementation, since the application is constantly updating its version, a new symbol table file will be generated for each version update, so the usage frequency of the old version of the symbol table file will continue to decrease. Therefore, the multi-level storage position of the first symbol table file can be adjusted by heat. As the usage heat of the symbol table file increases and decreases, the medium storing the symbol table file is dynamically adjusted. Specifically, the usage heat of the first symbol table file is determined; then the storage location of the first symbol table file is adjusted according to the usage heat, wherein the storage location includes a local disk, a network file system, and a cloud object storage, and the usage heat corresponding to the local disk, the network file system, and the cloud object storage decreases in turn.
[0163] Specifically, usage heat may refer to the number of times or frequency that a symbol table file is accessed within a fixed time period. All first symbol table files may be sorted in descending order according to usage heat, and the first symbol table files within the first heat range in the sorting result may be stored in a local disk, the first symbol table files within the second heat range in the sorting result may be stored in a network file system, and the first symbol table files within the third heat range in the sorting result may be stored in a cloud object storage, wherein the lower limit value of the first heat range is higher than the upper limit value of the second heat range, and the lower limit value of the second heat range is higher than the upper limit value of the third heat range. For example, the first heat range refers to 1% to 20% of the sorting result, the second heat range refers to 21% to 50% of the sorting result, and the third heat range refers to 51% to 100% of the sorting result. It should be noted that the usage heat of all first symbol table files is periodically determined, and the storage location of the first symbol table file is adjusted in time to improve resource utilization. For example, in the previous heat detection cycle, the usage heat ranking of the first symbol table file X is in the first heat range, that is, the first symbol table file X is stored in the local disk to facilitate frequent high-speed reading. In the current heat detection cycle, due to the version update of the application, the usage heat ranking of the first symbol table file X drops to the second heat range, then the first symbol table file X is removed from the local disk and transferred to the network file system for storage. With the version update of the application, the usage heat ranking of the first symbol table file X continues to decrease until it drops to the third heat range, then the first symbol table file X is deleted from the network file system and transferred to the cloud object storage for storage. Therefore, a multi-level storage architecture is used to achieve hot and cold data separation. By caching hot data, high-frequency requests can be quickly responded to, reducing network transmission, and regular cleaning of cold data helps to reduce memory usage.
[0164] In addition, the storage method of the first symbol table file can also adopt a cache backfill strategy, that is, after the first symbol table file is obtained by the low-speed storage, the first symbol table file can be cached to the high-speed storage (such as the local disk) to speed up subsequent access. Fig.10 , Fig.10 A schematic diagram of the first symbol table file reading process provided for an embodiment of the present disclosure, after receiving an instruction translation request, extracting a target address and a first symbol table file identifier from the instruction translation request, first searching in the local cache according to the target address whether there is an original instruction translation result corresponding to the target address, if the local cache has an original instruction translation result corresponding to the target address, then directly returning the original instruction translation result; if the local cache does not have an original instruction translation result corresponding to the target address, then searching in the local disk for the corresponding first symbol table file according to the first symbol table file identifier.
[0165] If the first symbol table file corresponding to the first symbol table file identifier exists in the local disk, the first symbol table file is called to perform instruction translation on the target address to obtain the original instruction translation result corresponding to the target address, and the original instruction translation result corresponding to the target address is written into the local cache, and then the original instruction translation result is returned; if the first symbol table file corresponding to the first symbol table file identifier does not exist in the local disk, the network file system is searched according to the first symbol table file identifier.
[0166] If there is a first symbol table file corresponding to the first symbol table file identifier in the network file system, the first symbol table file is called to perform instruction translation on the target address to obtain the original instruction translation result corresponding to the target address, and the original instruction translation result corresponding to the target address is written into the local cache, and then the original instruction translation result is returned; if there is no first symbol table file corresponding to the first symbol table file identifier in the network file system, it is searched in the cloud object storage according to the first symbol table file identifier.
[0167] If there is a first symbol table file corresponding to the first symbol table file identifier in the cloud object storage, the first symbol table file is downloaded from the cloud storage object and saved to the local disk, and then the first symbol table file is used to translate the instruction of the target address to obtain the original instruction translation result corresponding to the target address, and the original instruction translation result corresponding to the target address is written into the local cache, and then the original instruction translation result is returned.
[0168] Among them, if the first symbol table file is read in the network file system, the first symbol table file can be backed up to the local disk to speed up file access. Therefore, that is to say, after the first symbol table file is obtained by the low-speed storage (such as the network file system and the cloud object storage), the first symbol table file can be cached to the high-speed storage (such as the local disk) to speed up subsequent access, that is, to increase the use popularity of the first symbol table file within a period of time and change the storage medium of the first symbol table file. In addition, the local disk and the network file system will regularly remove the symbol table files with low use popularity.
[0169] Among them, after the developer uploads the first symbol table file to the server, the server can store the first symbol table file in the cloud object storage and save the meta information of the first symbol table file (save the storage path of the first symbol table file constructed by the product information in the cloud object storage). When instruction translation is required, according to the received instruction translation request, the service object identifier and the first symbol table file identifier are extracted, the meta information of the first symbol table file is found, and the first symbol table file is located through the storage path in the meta information, and then the first symbol table file is downloaded from the cloud object storage to the local disk for translation. It should be noted that the first symbol table file can be uploaded to the OLAP database for storage.
[0170] In one possible implementation, since the translation result of the original instruction usually contains function characters, class characters and line numbers, it occupies a small volume, and the target address in the same crash log may be translated multiple times, and the translation result of the original instruction is related to the current session or user, and has low long-term storage value, after obtaining the translation result of the original instruction corresponding to the target address, it can be saved in the local cache for quick access, and the expired cache can be automatically cleaned up through the least recently used (Least Recently Used, LRU) strategy. For example, when user A analyzes the crash log, the target address 0x10388 is translated to obtain the corresponding translation result of the original instruction (containing function characters and line numbers), and then this translation result is saved in the local cache. When user B analyzes the same log, the local cache can be directly hit without reloading the first symbol table file. Fig.11 , Fig.11 The effect diagram of the data hierarchical cache provided for the embodiment of the present disclosure shows that the hit rate of the translation result in the local cache is as high as 77%, that is, after receiving the instruction translation request, there is a 77% probability that the corresponding translation result can be directly found in the local cache, and the remaining 23% probability requires calling the first symbol table file for translation, wherein, after calling the first symbol table file to translate the target address, the translation result will be saved in the local cache, so that subsequent requests for the same address can directly hit the local cache. This is because the target addresses in the crash log are often concentrated in a few core modules (such as the main business logic, commonly used library functions), and the symbol translation requirements of these modules are frequent. Therefore, when these translation results are saved in the local cache, they can cover and respond to most of the subsequent requests, reduce repeated calculations, and adopt the LRU strategy in the local cache, give priority to retaining the most recently accessed data, and automatically eliminate cold data that has not been used for a long time, further improving the hit rate of the translation result in the local cache.
[0171] When obtaining the first symbol table file in multi-level storage, the hit rate in the local disk is as high as 76.5%, the hit rate in the network file system is 23%, and the hit rate in the cloud object storage is only 0.5%. This is because after the new first symbol table file is made, the new first symbol table file will be uploaded to the cloud object storage, and after the first call to the new first symbol table file, the first symbol table file will be stored in the local disk. At the same time, the network file system will transfer the first symbol table file with high frequency access in a time window to the local disk for storage, and both the local disk and the network file system will use the LRU strategy to transfer the first symbol table file with low frequency access in a time window. The first symbol table file is transferred to the cloud object storage, so the translation result is cached locally to reduce repeated translation of the same request, and the first symbol table file with high frequency of access is stored in the local disk, so that the computing performance can grow linearly with the horizontal expansion. Since the storage capacity of the network file system is larger than the storage capacity of the local disk, and the access speed is faster than the access speed of the cloud object storage, the first symbol table file with medium access frequency is stored in the network file system, reducing the frequency of downloading the first symbol table file from the cloud object storage, and the cloud object storage can compress and store each first symbol table file, and can store a large number of first symbol table files.
[0172] In one possible implementation, when an instruction translation request carries an instruction to be translated, the instruction translation request can be parsed to obtain the instruction to be translated and the second symbol table file identifier, wherein the instruction to be translated includes the character to be translated and the original instruction characters in the original instruction except the second function character, and the character to be translated is the character before translation corresponding to the second function character. Then, the second symbol table file is obtained according to the second symbol table file identifier, and the second row number of the character to be translated in the second symbol dictionary is retrieved from the second symbol information of the second symbol table file according to the character to be translated, wherein the second symbol information is used to store the mapping relationship between the character to be translated and the second row number; then, the second symbol dictionary is obtained according to the instruction to be translated, and the second function character and the third row number of the original instruction are obtained in the second symbol dictionary according to the second row number as the translation result of the instruction to be translated. The above process is applicable to the scenario of translating character instructions, and supports mapping and translation of multiple languages or custom symbols to meet different usage scenarios.
[0173] Among them, the character to be translated can refer to an identifier in the string resource, and the character that needs to be translated by calling the second symbol table file to identify the corresponding symbol table file can be obtained. After translation, the second function character in the original instruction can be obtained. The second function character can refer to a string representing a function name in a running code of an application, and the instruction to be translated refers to the other part including the translated characters and the original instruction, namely the original instruction character. The original instruction character means the character part that does not need to be translated, that is, a part of the instruction in the actual running code, including at least one of the file character and the class character. For example, assuming that the instruction to be translated is (class character: character to be translated), the translation result obtained after translation is (class character: second function character: third line number), and the third line number can refer to the line number of another original instruction in the running code of the application, which is used to identify different original instructions.
[0174] Among them, the second symbol table file identifier can refer to the unique identifier of the symbol table file used to translate the characters to be translated. The storage method and acquisition method of the second symbol table file can refer to the storage method and acquisition method of the first symbol table file proposed in the above embodiment, and the details are not repeated here. The second symbol table file stores second symbol information, and the second symbol information is used to store the mapping relationship between the character to be translated and the second row number. The second symbol dictionary may refer to a data structure for storing the mapping relationship between the second row number and the function character, and is used to map the function character to the unique identifier, i.e., the second row number. For example, the character to be translated and the corresponding row number may be written as an entry in the second symbol information, and the second row number is used to locate the corresponding function character in the second symbol dictionary. Therefore, the character to be translated may be used as a key to find the corresponding entry in the second symbol information, and the second row number corresponding to the character to be translated may be determined. Then, the second row number may be used to locate the function character and row number of the corresponding row in the second symbol dictionary, and the function character may be used as the second function character, and the row number may be used as the third row number where the original instruction is located. Then, the second function character and the third row number may be combined as the translation result of the character to be translated, or, the second function character and the third row number may replace the character to be translated in the instruction to be translated, and the translation result of the instruction to be translated may be obtained.
[0175] In one possible implementation, the instruction to be translated contains the row number of the original instruction. Therefore, after translating the characters to be translated using the second symbol table file and obtaining the second row number in the second symbol dictionary, the second function character can be obtained in the second symbol dictionary according to the second row number, without further determining the third row number of the original instruction in the second symbol dictionary.
[0176] In one possible implementation, the process of creating the second candidate dictionary may be to first obtain the service object identifier of each service object, the second original dictionary of each service object, and the second symbol table file identifier, obtain multiple second function characters and the second row number corresponding to each second function character in the second original dictionary; then, create a second candidate dictionary, write the second function characters and the corresponding second row numbers as entries to the second candidate dictionary in sequence, until the storage space occupied by the second candidate dictionary reaches a preset threshold, create the next second candidate dictionary to continue writing entries, until all the second function characters and the corresponding second row numbers are written, and assign corresponding instruction indexes to each second candidate dictionary; then, the second symbol table file identifier, the service object identifier, the multiple instruction indexes and the multiple second candidate dictionaries are associated and stored, and by splitting the second original dictionary, limiting the size of the dictionary fragments helps to manage the storage space, and the use of associated storage and instruction indexes can increase the speed of data search and access, and improve the efficiency of sorting and translation.
[0177] Among them, the instruction index can obtain the second function character corresponding to the first entry written in the second candidate dictionary as the index character, or the instruction index can obtain the second function character corresponding to the first entry written in the second candidate dictionary and the second function character corresponding to the last entry as the index character, so that the instruction index can be used to indicate the mapping range of the second candidate dictionary.
[0178] Among them, the content of the second candidate dictionary is different from that of the first candidate dictionary. The second candidate dictionary chooses to write function characters and row numbers as entries to construct a mapping relationship between function characters and row numbers, while the first candidate dictionary chooses to write function characters and function numbers as entries to construct a mapping relationship between function characters and function numbers.
[0179] The service object uploads the second symbol table file and the second candidate dictionary to the cloud object storage for associated storage. Specifically, refer to Figure 7 , the service object identifier is used as the first-level directory, the second symbol table file identifier is used as the second-level directory, and the second candidate dictionary and the corresponding instruction index are associated as the third-level directory under the corresponding second-level directory.
[0180] In a possible implementation, the second candidate dictionary can be obtained by splitting the second original dictionary, and the fragmented second candidate dictionary, instruction index and second symbol table file identifier are associated and stored, so that multiple second candidate dictionaries stored in fragments can be determined according to the second symbol table file identifier, and the character to be translated is matched with the instruction index corresponding to each second candidate dictionary; when there is no instruction index matching the character to be translated, each index character is sorted in a preset order to obtain an index character sequence; the first character and the second character adjacent to the character to be translated in the index character sequence are determined, wherein the ranking of the first character in the index character sequence is less than the ranking of the second character in the index character sequence; the second candidate dictionary corresponding to the instruction index where the first character is located is obtained as the second symbol dictionary, and each index character can be used for direct indexing, and each index character can be used to construct an index character sequence to achieve a larger index range. This flexible indexing method not only has good scalability and adaptability, but also does not require the configuration of a complete index range in the instruction index, reduces the data storage of the instruction index, and saves storage space.
[0181] Reference Fig.12 , Fig.12 A schematic diagram of the process of obtaining a second symbol dictionary provided for an embodiment of the present disclosure, assuming that a character to be translated is M, three second candidate dictionaries associated with the storage can be located in the cloud object storage through the second symbol table file identifier, namely, second candidate dictionary 1, second candidate dictionary 2, and second candidate dictionary 3, instruction index 1 of second candidate dictionary 1 contains index character A, instruction index 2 of second candidate dictionary 2 contains index character G, and instruction index 3 of second candidate dictionary 3 contains index character P, and the character M to be translated is matched with these instruction indexes. Obviously, these instruction indexes do not have the same index character as the character M to be translated, therefore, the index characters of all instruction indexes are sorted in a preset order to obtain an index character sequence, which is arranged from small to large in ranking, namely, index character A, index character G, index character P, wherein it is found that the character M to be translated is located between index character G and index character P in the index character sequence, and it is determined that index character G is the first character, so that the second candidate dictionary 2 corresponding to instruction index 2 where index character G is located is used as the second symbol dictionary.
[0182] Reference Fig.13 , Fig.13This is a schematic diagram of the instruction translation method in the related technology. When an exception occurs in the client application, the client SDK collects relevant information, generates a crash log, and reports the instruction translation request to the server. After receiving the instruction translation request, the server parses the instruction translation request, extracts the target address and the symbol table identifier, and checks whether the target address hits the local cache. If so, the result is returned directly. If not, the corresponding symbol table is searched in the local cache according to the symbol table identifier. If so, the symbol table is directly used to symbolically translate the target address. If not, the full symbol table is downloaded from the distributed file system in real time to the local cache, and the symbol table is used in the local cache to symbolically translate the target address. After translation, the result is returned to the client. It can be seen that the instruction translation method used in the related art requires real-time downloading of the full symbol table, which is very large and takes a long time. In addition, parallel downloading of files under high concurrency will increase the number of read and write times per second of the distributed file system, which can easily cause the distributed file system to crash. At the same time, the related art usually uses the Address To Symbol (ATOS) tool to read the symbol table and then translate it. The ATOS tool is a command line tool that is mainly used to convert memory addresses into readable function names, source code file names and line numbers. However, when using the ATOS tool for translation, each address can only be translated one by one. At the same time, the ATOS tool relies on a complete symbol table. Even translating a memory address requires a full symbol table, which results in low efficiency in instruction translation in the related art.
[0183] In one possible implementation, referring to Fig.14 , Fig.14 The effect diagram of instruction translation provided by the embodiment of the present disclosure is as follows: when a terminal crashes while running a client application, the client SDK will collect and organize the data at the time of the crash, generate a crash log, and report an instruction translation request to the symbol translation server in the hope of translating the original instructions in the crash log. After receiving the instruction translation request, the symbol translation server parses the instruction translation request to obtain the target address and the first symbol table file identifier, such as Fig.14As shown in , the first instruction translation request reported by the client carries multiple target addresses to be translated, and then searches the server local cache for a matching translation result based on the target address. If so, the translation result is directly returned to the terminal. Otherwise, the first symbol table file is sequentially obtained from the server local disk, network file system, and cloud object storage according to the first symbol table file identifier until the first symbol table file is successfully obtained; the target address is then translated using the first symbol table file to obtain the function number corresponding to the target address and the first line number of the original instruction, and then the first symbol dictionary is used to retrieve the first function character corresponding to the function number, the first function character is combined with the first line number to obtain the translation result of the original instruction, the translation result is associated with the target address and written into the server local cache, and then the translation result is returned to the terminal, as shown in FIG. Fig.14 As shown, the first instruction translation request carrying multiple target addresses to be translated is restored into function characters and line numbers.
[0184] When the instruction translation request carries the instruction to be translated, the instruction to be translated and the second symbol table file identifier are extracted from the instruction translation request, such as Fig.14 As shown, the instruction to be translated in the second instruction translation request contains the original instruction character and the character to be translated; similarly, according to the second symbol table file identifier, the second symbol table file is sequentially obtained in the server local disk, the network file system and the cloud object storage until the second symbol table file is successfully obtained, and then the second symbol table file is used to translate the character to be translated to obtain the second row number, and then the second symbol dictionary is used to retrieve the second function character corresponding to the second row number, the original instruction character is combined with the second function character to obtain the translation result of the original instruction, the translation result is associated with the target address and written into the server local cache, and then the translation result is returned to the terminal, as shown in FIG. Fig.14 As shown, the characters to be translated carried in the second instruction translation request are restored to the second function characters.
[0185] It is worth noting that when the target address is translated using the first symbol table file, since the first symbol table file is reconstructed, each node in the tree data structure manages an address range of a fixed size. Therefore, when searching for a node index that matches the target address in the tree data structure, the index of the next layer of nodes can be obtained by direct address calculation, without relying on child node pointers for positioning or relying on the linked list structure of leaf nodes for range scanning. The first symbol information is stored using function numbers instead of complete function characters, and the tree data structure and the first symbol information are used to complete the establishment of a mapping relationship between the target address, the function number and the first row number. The function number of the target address and the first row number of the original instruction are found through the tree data structure, and the first symbol dictionary obtained by the target address is used to retrieve the first function character corresponding to the function number. The first function character and the first line number are combined to restore the translation result of the original instruction, which not only effectively reduces the size of the symbol table file, but also effectively reduces the time complexity of the query stage. Specifically, for example, when translating an instruction translation request, it takes about 17 seconds to download the original symbol table file (with a size of 2.4GB), while when using the instruction translation method provided in the embodiment of the present invention to translate the same instruction translation request, it takes about 5 seconds to download the first symbol table file (with a size of 700MB). It can be seen that by using the instruction translation method provided in the embodiment of the present invention, the symbol table file required for the same instruction translation request can be optimized, and the original symbol table file with a size of 2.4GB can be trimmed to obtain a first symbol table file with a size of 700MB, and the optimized size is compressed to 30% of the original file size, thereby reducing download time. Similarly, after adopting the instruction translation method provided by the embodiment of the present disclosure, compared with the binary search method adopted in the related art, the time complexity can be reduced from O(logn) to O(1), and the average response time for the same instruction translation request is reduced from the original average of 2000 milliseconds to 1224 milliseconds, and the average throughput is increased from the original 100 to 142.
[0186] The instruction translation method provided by the embodiment of the present disclosure is described in detail below with reference to a specific example.
[0187] Reference Fig.15 As shown, Fig.15 This is a specific flow chart of the instruction translation method provided by a specific example. Fig.15 In the example, the instruction translation method may include the following steps 1501 to 1516.
[0188] Step 1501: Receive an instruction translation request. When the instruction translation request carries a target address of an original instruction, obtain the target address and a first symbol table file identifier in the instruction translation request.
[0189] Step 1502: Obtain a first symbol table file from a local disk according to the first symbol table file identifier.
[0190] Step 1503: When the first symbol table file does not exist in the local disk, the first symbol table file is obtained in the network file system according to the first symbol table file identifier.
[0191] Step 1504: When the first symbol table file does not exist in the network file system, the first symbol table file is obtained in the cloud object storage according to the first symbol table file identifier.
[0192] Step 1505: Retrieve a first reference address that matches the target address from the root node of the tree data structure of the first symbol table file.
[0193] Step 1506: Determine the element position of the first reference address in the root node.
[0194] Step 1507: Obtain the system identifier of the operating system from which the instruction translation request comes from in the instruction translation request, and determine the memory block size corresponding to the operating system according to the system identifier.
[0195] Step 1508: Determine the node address offset based on the product of the element position and the memory block size.
[0196] Step 1509: Obtain the node start address of the leaf node associated with the root node, and determine the destination node address according to the node start address and the node address offset.
[0197] Step 1510: Retrieve a second reference address matching the target address from a branch node corresponding to the destination node address.
[0198] Step 1511: Determine the leaf node associated with the second reference address as the target node, and read the symbol offset from the target node.
[0199] Step 1512: Retrieve the function number of the target address and the first line number where the original instruction is located from the first symbol information of the first symbol table file according to the symbol offset.
[0200] Step 1513: Obtain the tag file, and retrieve the symbol dictionary index in the tag file according to the target address.
[0201] Step 1514: Obtain a first symbol dictionary from a plurality of first candidate dictionaries stored in the shards according to the first symbol table file identifier and the symbol dictionary index.
[0202] Step 1515: Retrieve the first function character of the original instruction in the first symbol dictionary according to the function number.
[0203] Step 1516: Obtain the translation result of the original instruction based on the first function character and the first line number.
[0204] Based on this, by receiving an instruction translation request, when the instruction translation request carries the target address of the original instruction, the target address and the first symbol table file identifier are obtained in the instruction translation request, the first symbol table file is obtained according to the first symbol table file identifier, the target node is retrieved from the tree data structure of the first symbol table file according to the target address, the symbol offset is read from the target node, and the function number of the target address and the first line number of the original instruction are retrieved from the first symbol information of the first symbol table file according to the symbol offset. Since the first symbol table file contains the tree data structure and the first symbol information, the content is reconstructed compared to the symbol table file in the related technology, and a mapping relationship can be established between the target address, the function number and the first line number based on the tree data structure and the first symbol information, thereby reducing the file size. On this basis, the first symbol dictionary is obtained according to the target address, the first function character of the original instruction is retrieved from the first symbol dictionary according to the function number, and the translation result of the original instruction is obtained based on the first function character and the first line number. Compared with the binary search method, the above translation process optimizes the retrieval process, can reduce the time complexity of instruction translation, and improve the instruction translation efficiency.
[0205] It is to be understood that, although the steps in the above-mentioned various flow charts are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in the present embodiment, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the above-mentioned flow charts can include a plurality of steps or a plurality of stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of the steps or stages in other steps or other steps.
[0206] Reference Fig.16 , Fig.16 An optional structural diagram of an instruction translation device provided in an embodiment of the present disclosure, the instruction translation device 1600 includes:
[0207] The request processing module 1610 is used to receive an instruction translation request, and when the instruction translation request carries a target address of an original instruction, obtain the target address from the instruction translation request;
[0208] The tree data structure processing module 1620 is used to obtain the first symbol table file, retrieve the target node in the tree data structure of the first symbol table file according to the target address, and read the symbol offset from the target node;
[0209] A symbol information processing module 1630 is used to retrieve the function number of the target address and the first line number of the original instruction from the first symbol information of the first symbol table file according to the symbol offset, wherein the first symbol information is used to store the mapping relationship between the symbol offset, the function number and the first line number;
[0210] The symbol dictionary processing module 1640 is used to obtain a first symbol dictionary according to the target address, retrieve the first function character of the original instruction in the first symbol dictionary, and obtain the translation result of the original instruction based on the first function character and the first line number.
[0211] In a possible implementation, the tree data structure processing module 1620 is further used to:
[0212] Retrieving a first reference address matching the target address from a root node of a tree data structure of a first symbol table file, wherein the first reference address is a starting address of an address range where the target address is located;
[0213] Retrieving a second reference address matching the target address from a branch node associated with the first reference address;
[0214] The leaf node associated with the second reference address is determined as the target node.
[0215] In a possible implementation, the tree data structure processing module 1620 is further used to:
[0216] Determine the element position of the first reference address in the root node, and determine the node address offset according to the element position;
[0217] Get the node start address of the leaf node associated with the root node, and determine the destination node address based on the node start address and the node address offset;
[0218] A second reference address matching the target address is retrieved from a branch node corresponding to the destination node address.
[0219] In a possible implementation, the tree data structure processing module 1620 is further used to:
[0220] Obtaining a system identifier of an operating system from which the instruction translation request comes in the instruction translation request, and determining a memory block size corresponding to the operating system according to the system identifier;
[0221] The node address offset is determined by the product of the element position and the memory block size.
[0222] In a possible implementation, the symbol dictionary processing module 1640 is further used to:
[0223] Obtain a tag file, and retrieve a symbol dictionary index in the tag file according to the target address, wherein the symbol dictionary index is used to indicate an address range where the target address is located;
[0224] A first symbol dictionary is obtained from a plurality of first candidate dictionaries stored in the shards according to the first symbol table file identifier and the symbol dictionary index.
[0225] In a possible implementation, the symbol dictionary processing module 1640 is further used to:
[0226] Obtaining a service object identifier of each service object, a first original dictionary of each service object, and a first symbol table file identifier;
[0227] Divide the first original dictionary into a plurality of first candidate dictionaries, and assign a corresponding symbol dictionary index to each first candidate dictionary;
[0228] The tag file, the first symbol table file identifier, the service object identifier, multiple symbol dictionary indexes, and multiple first candidate dictionaries are stored in association.
[0229] In a possible implementation, the tree data structure processing module 1620 is further used to:
[0230] Obtaining a first symbol table file from a local disk according to the first symbol table file identifier;
[0231] When the first symbol table file does not exist in the local disk, obtaining the first symbol table file in the network file system according to the first symbol table file identifier;
[0232] When the first symbol table file does not exist in the network file system, the first symbol table file is obtained in the cloud object storage according to the first symbol table file identifier.
[0233] In a possible implementation, the tree data structure processing module 1620 is further used to:
[0234] Determine the usage popularity of the first symbol table file;
[0235] The storage location of the first symbol table file is adjusted according to the usage heat, wherein the storage location includes a local disk, a network file system, and a cloud object storage, and the usage heat corresponding to the local disk, the network file system, and the cloud object storage decreases in sequence.
[0236] In a possible implementation, the instruction translation device 1600 further includes a symbol table file making module, and the symbol table file making module is used to:
[0237] Obtain an original symbol table file, and extract multiple target addresses, first function characters corresponding to each target address, and first row numbers corresponding to each first function character from the original symbol table file;
[0238] Creating a tree data structure based on multiple target addresses and symbol offsets, assigning a function number to a first function character, creating first symbol information based on the symbol offset, the function number and the first row number, and encapsulating the tree data structure and the first symbol information into a first symbol table file;
[0239] A first symbol dictionary is created based on the function number and the first function character.
[0240] In one possible implementation,
[0241] The request processing module 1610 is also used to obtain the instruction to be translated and the second symbol table file identifier in the instruction translation request when the instruction translation request carries the instruction to be translated, wherein the instruction to be translated includes the character to be translated and the original instruction character except the second function character in the original instruction, and the character to be translated is the character before translation corresponding to the second function character;
[0242] The tree data structure processing module 1620 is further used to obtain the second symbol table file according to the second symbol table file identifier, and retrieve the second row number of the character to be translated in the second symbol dictionary according to the second symbol information of the character to be translated, wherein the second symbol information is used to store the mapping relationship between the character to be translated and the second row number;
[0243] The symbol dictionary processing module 1640 is further used to obtain a second symbol dictionary according to the instruction to be translated, and obtain the second function character and the third row number of the original instruction in the second symbol dictionary according to the second row number as the translation result of the instruction to be translated.
[0244] In a possible implementation, the symbol dictionary processing module 1640 is further used to:
[0245] Determine multiple second candidate dictionaries stored in the fragments according to the second symbol table file identifier, and match the instruction to be translated with the instruction index corresponding to each second candidate dictionary, wherein the instruction index includes an index character and an original instruction character;
[0246] When there is no instruction index matching the instruction to be translated, sorting the index characters in a preset order to obtain an index character sequence;
[0247] Determine a first character and a second character that are adjacent to the character to be translated in the index character sequence, wherein the ranking of the first character in the index character sequence is lower than the ranking of the second character in the index character sequence;
[0248] A second candidate dictionary corresponding to the instruction index where the first character is located is obtained as the second symbol dictionary.
[0249] In a possible implementation, the symbol dictionary processing module 1640 is further used to:
[0250] Obtaining a service object identifier of each service object, a second original dictionary of each service object, and a second symbol table file identifier, and obtaining a plurality of second function characters and a second row number corresponding to each second function character in the second original dictionary;
[0251] Creating a second candidate dictionary, writing the second function characters and the corresponding second row numbers as entries into the second candidate dictionary in sequence, until the storage space occupied by the second candidate dictionary reaches a preset threshold, creating the next second candidate dictionary to continue writing entries, until all the second function characters and the corresponding second row numbers are written, and assigning a corresponding instruction index to each second candidate dictionary;
[0252] The second symbol table file identifier, the service object identifier, the plurality of instruction indexes, and the plurality of second candidate dictionaries are stored in association with each other.
[0253] The above-mentioned instruction translation device 1600 and the instruction translation method are based on the same inventive concept, by receiving an instruction translation request, when the instruction translation request carries the target address of the original instruction, the target address is obtained in the instruction translation request, the first symbol table file is obtained, the target node is retrieved from the tree data structure of the first symbol table file according to the target address, the symbol offset is read from the target node, and the function number of the target address and the first line number of the original instruction are retrieved from the first symbol information of the first symbol table file according to the symbol offset. Since the first symbol table file contains the tree data structure and the first symbol information, the content is reconstructed compared with the symbol table file in the related art, and the target address, function number and first line number can be mapped based on the tree data structure and the first symbol information to achieve file size reduction. On this basis, the first symbol dictionary is obtained according to the target address, the first function character of the original instruction is retrieved from the first symbol dictionary according to the function number, and the translation result of the original instruction is obtained based on the first function character and the first line number. Compared with the binary search method, the above-mentioned translation process optimizes the retrieval process, which can reduce the time complexity of instruction translation and improve the instruction translation efficiency.
[0254] The electronic device for executing the above instruction translation method provided in the embodiment of the present disclosure may be a terminal. Fig.17 , Fig.17This is a partial structural block diagram of a terminal provided in an embodiment of the present disclosure, and the terminal includes: a camera assembly 1710, a first memory 1720, an input unit 1730, a display unit 1740, a sensor 1750, an audio circuit 1760, a wireless fidelity (WiFi) module 1770, a first processor 1780, and a first power supply 1790. Those skilled in the art can understand that Fig.17 The terminal structure shown in the figure does not constitute a limitation on the terminal, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0255] The camera assembly 1710 can be used to capture images or videos. Optionally, the camera assembly 1710 includes a front camera and a rear camera. Typically, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize the panoramic shooting and the VR (Virtual Reality) shooting function or other fusion shooting functions.
[0256] The first memory 1720 may be used to store software programs and modules. The first processor 1780 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the first memory 1720 .
[0257] The input unit 1730 may be used to receive input digital or character information and generate key signal input related to the terminal's settings and function control. Specifically, the input unit 1730 may include a touch panel 1731 and other input devices 1732 .
[0258] The display unit 1740 may be used to display input information or provided information and various menus of the terminal. The display unit 1740 may include a display panel 1741 .
[0259] The audio circuit 1760 , the speaker 1761 , and the microphone 1762 may provide an audio interface.
[0260] The first power source 1790 may be alternating current, direct current, a disposable battery, or a rechargeable battery.
[0261] The number of sensors 1750 may be one or more, and the one or more sensors 1750 include but are not limited to: acceleration sensors, gyroscope sensors, pressure sensors, optical sensors, etc. Among them:
[0262] The acceleration sensor can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal. For example, the acceleration sensor can be used to detect the components of gravity acceleration on the three coordinate axes. The first processor 1780 can control the display unit 1740 to display the user interface in a horizontal view or a vertical view according to the gravity acceleration signal collected by the acceleration sensor. The acceleration sensor can also be used for collecting game or user motion data.
[0263] The gyroscope sensor can detect the body direction and rotation angle of the terminal, and the gyroscope sensor can cooperate with the acceleration sensor to collect the user's 3D actions on the terminal. The first processor 1780 can implement the following functions based on the data collected by the gyroscope sensor: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0264] The pressure sensor can be set in the side frame of the terminal and / or the lower layer of the display unit 1740. When the pressure sensor is set in the side frame of the terminal, the user's holding signal of the terminal can be detected, and the first processor 1780 performs left and right hand recognition or shortcut operation according to the holding signal collected by the pressure sensor. When the pressure sensor is set in the lower layer of the display unit 1740, the first processor 1780 controls the operability controls on the UI interface according to the user's pressure operation on the display unit 1740. The operability control includes at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0265] The optical sensor is used to collect the ambient light intensity. In one embodiment, the first processor 1780 can control the display brightness of the display unit 1740 according to the ambient light intensity collected by the optical sensor. Specifically, when the ambient light intensity is high, the display brightness of the display unit 1740 is increased; when the ambient light intensity is low, the display brightness of the display unit 1740 is decreased. In another embodiment, the first processor 1780 can also dynamically adjust the shooting parameters of the camera assembly 1710 according to the ambient light intensity collected by the optical sensor.
[0266] In this embodiment, the first processor 1780 included in the terminal can execute the instruction translation method of the previous embodiment.
[0267] The electronic device for executing the above instruction translation method provided in the embodiment of the present disclosure may also be a server. Fig.18 , Fig.18Partial structural block diagram of a server provided in an embodiment of the present disclosure. The server may have relatively large differences due to different configurations or performances, and may include one or more second processors 1810 and a second memory 1830, and one or more storage media 1840 (e.g., one or more mass storage devices) storing application programs 1843 or data 1842. Among them, the second memory 1830 and the storage medium 1840 may be short-term storage or persistent storage. The program stored in the storage medium 1840 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the second processor 1810 may be configured to communicate with the storage medium 1840 to execute a series of instruction operations in the storage medium 1840 on the server.
[0268] The server may also include one or more second power supplies 1820, one or more wired or wireless network interfaces 1850, one or more input and output interfaces 1860, and / or one or more operating systems 1841, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0269] The second processor 2010 in the server may be configured to execute the instruction translation method.
[0270] The embodiments of the present disclosure further provide a computer-readable storage medium, which is used to store a computer program, and the computer program is used to execute the instruction translation method of each of the aforementioned embodiments.
[0271] The embodiment of the present disclosure also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above-mentioned instruction translation method.
[0272] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate to describe the embodiments of the present disclosure, such as being able to be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0273] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0274] It should be understood that in the description of the embodiments of the present disclosure, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood to not include the number, and above, below, within, etc. are understood to include the number.
[0275] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0276] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0277] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0278] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store program codes.
[0279] It should also be understood that the various implementations provided in the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.
[0280] The above is a specific description of the preferred implementation of the present disclosure, but the present disclosure is not limited to the above-mentioned implementation mode. Technical personnel familiar with the field can also make various equivalent deformations or substitutions under the shared conditions without violating the spirit of the present disclosure. These equivalent deformations or substitutions are all included in the scope defined by the claims of the present disclosure.
Claims
1. A method for translating instructions, characterized in that: include: receiving an instruction translation request, and when the instruction translation request carries a target address of an original instruction, obtaining the target address from the instruction translation request; Obtaining a first symbol table file, retrieving a target node in a tree data structure of the first symbol table file according to the target address, and reading a symbol offset from the target node; Retrieving the function number of the target address and the first row number where the original instruction is located in the first symbol information of the first symbol table file according to the symbol offset, wherein the first symbol information is used to store a mapping relationship between the symbol offset, the function number and the first row number; A first symbol dictionary is obtained according to the target address, a first function character of the original instruction is retrieved from the first symbol dictionary according to the function number, and a translation result of the original instruction is obtained based on the first function character and the first line number.
2. The instruction translation method according to claim 1, characterized in that: The step of retrieving a target node from the tree data structure of the first symbol table file according to the target address comprises: Retrieving a first reference address matching the target address from a root node of the tree data structure of the first symbol table file, wherein the first reference address is a starting address of an address range where the target address is located; Retrieving a second reference address matching the target address from a branch node associated with the first reference address; The leaf node associated with the second reference address is determined as the target node.
3. The instruction translation method according to claim 2, characterized in that: The step of retrieving a second reference address matching the target address from a branch node associated with the first reference address includes: Determine an element position of the first reference address in the root node, and determine a node address offset according to the element position; Obtaining a node start address of a leaf node associated with the root node, and determining a destination node address according to the node start address and the node address offset; A second reference address matching the target address is retrieved from a branch node corresponding to the destination node address.
4. The instruction translation method according to claim 3, characterized in that: Determining the node address offset according to the element position includes: Obtaining, in the instruction translation request, a system identifier of an operating system from which the instruction translation request comes, and determining a memory block size corresponding to the operating system according to the system identifier; The node address offset is determined according to the product of the element position and the memory block size.
5. The instruction translation method according to claim 1, characterized in that: The instruction translation request also carries a first symbol table file identifier of the first symbol table file, and acquiring a first symbol dictionary according to the target address includes: Obtain a tag file, and retrieve a symbol dictionary index in the tag file according to the target address, wherein the symbol dictionary index is used to indicate an address range where the target address is located; A first symbol dictionary is obtained from a plurality of first candidate dictionaries stored in the shards according to the first symbol table file identifier and the symbol dictionary index.
6. The instruction translation method according to claim 5, characterized in that: Before acquiring the first symbol dictionary from a plurality of first candidate dictionaries stored in the shards according to the first symbol table file identifier and the symbol dictionary index, the instruction translation method further includes: Obtaining a service object identifier of each service object, a first original dictionary of each service object, and a first symbol table file identifier; Divide the first original dictionary into a plurality of the first candidate dictionaries, and assign a corresponding symbol dictionary index to each of the first candidate dictionaries; The tag file, the first symbol table file identifier, the service object identifier, a plurality of the symbol dictionary indexes, and a plurality of the first candidate dictionaries are stored in association.
7. The instruction translation method according to claim 1, characterized in that: The instruction translation request also carries a first symbol table file identifier of the first symbol table file, and obtaining the first symbol table file includes: Obtaining a first symbol table file in a local disk according to the first symbol table file identifier; When the first symbol table file does not exist in the local disk, obtaining the first symbol table file in the network file system according to the first symbol table file identifier; When the first symbol table file does not exist in the network file system, the first symbol table file is obtained in the cloud object storage according to the first symbol table file identifier.
8. The instruction translation method according to claim 7, characterized in that: The instruction translation method further includes: Determining the usage popularity of the first symbol table file; The storage location of the first symbol table file is adjusted according to the usage heat, wherein the storage location includes the local disk, the network file system and the cloud object storage, and the usage heat corresponding to the local disk, the network file system and the cloud object storage decreases in sequence.
9. The instruction translation method according to claim 1, characterized in that: Before obtaining the first symbol table file, the instruction translation method further includes: Obtain an original symbol table file, and extract from the original symbol table file a plurality of the target addresses, the first function characters corresponding to the respective target addresses, and the first row numbers corresponding to the respective first function characters; Creating the tree data structure based on the plurality of target addresses and the symbol offsets, assigning the function number to the first function character, creating the first symbol information based on the symbol offset, the function number and the first row number, and encapsulating the tree data structure and the first symbol information into the first symbol table file; The first symbol dictionary is created based on the function number and the first function character.
10. The instruction translation method according to claim 1, characterized in that: The instruction translation method further includes: When the instruction translation request carries an instruction to be translated, the instruction to be translated and a second symbol table file identifier are obtained in the instruction translation request, wherein the instruction to be translated includes characters to be translated and original instruction characters in the original instruction except for the second function characters, and the characters to be translated are characters before translation corresponding to the second function characters; Acquire a second symbol table file according to the second symbol table file identifier, and retrieve the second row number of the character to be translated in the second symbol dictionary according to the character to be translated in the second symbol information of the second symbol table file, wherein the second symbol information is used to store a mapping relationship between the character to be translated and the second row number; The second symbol dictionary is obtained according to the character to be translated, and the second function character and the third row number where the original instruction is located are obtained from the second symbol dictionary according to the second row number as the translation result of the instruction to be translated.
11. The instruction translation method according to claim 10, characterized in that: The step of obtaining the second symbol dictionary according to the to-be-translated character comprises: Determine a plurality of second candidate dictionaries stored in the fragment according to the second symbol table file identifier, and match the characters to be translated with instruction indexes corresponding to the respective second candidate dictionaries, wherein the instruction indexes include index characters; When there is no instruction index matching the character to be translated, sorting each of the index characters in a preset order to obtain an index character sequence; Determine a first character and a second character adjacent to the character to be translated in the index character sequence, wherein the ranking of the first character in the index character sequence is lower than the ranking of the second character in the index character sequence; The second candidate dictionary corresponding to the instruction index where the first character is located is obtained as the second symbol dictionary.
12. The instruction translation method according to claim 11, characterized in that: Before determining the plurality of second candidate dictionaries for shard storage according to the second symbol table file identifier, the instruction translation method further includes: Obtaining a service object identifier of each service object, a second original dictionary of each service object, and a second symbol table file identifier, and obtaining a plurality of second function characters and the second row number corresponding to each second function character in the second original dictionary; creating a second candidate dictionary, writing the second function characters and the corresponding second row numbers as entries into the second candidate dictionary in sequence, until the storage space occupied by the second candidate dictionary reaches a preset threshold, creating the next second candidate dictionary to continue writing the entries, until all the second function characters and the corresponding second row numbers are written, and assigning the corresponding instruction index to each second candidate dictionary; The second symbol table file identifier, the service object identifier, a plurality of the instruction indexes, and a plurality of the second candidate dictionaries are stored in association with each other.
13. A command translation device, characterized in that: include: A request processing module, configured to receive an instruction translation request, and when the instruction translation request carries a target address of an original instruction, obtain the target address from the instruction translation request; A tree data structure processing module, used for obtaining a first symbol table file, retrieving a target node in the tree data structure of the first symbol table file according to the target address, and reading a symbol offset from the target node; A symbol information processing module, used for retrieving the function number of the target address and the first line number of the original instruction from the first symbol information of the first symbol table file according to the symbol offset, wherein the first symbol information is used for storing a mapping relationship between the symbol offset, the function number and the first line number; The symbol dictionary processing module is used to obtain a first symbol dictionary according to the target address, retrieve a first function character of the original instruction in the first symbol dictionary, and obtain a translation result of the original instruction based on the first function character and the first line number.
14. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the instruction translation method according to any one of claims 1 to 12 is implemented.
15. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the instruction translation method according to any one of claims 1 to 12 is implemented.
16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the instruction translation method according to any one of claims 1 to 12 is implemented.
Citation Information
Cited By
Binary translation method, translator, electronic equipment and readable storage medium
CN120353469A
Dynamic binary translation acceleration method and system, medium and product
CN120596103A
Redundant symbol searching method and device, terminal and medium
CN121501293A
Redundant symbol searching method and apparatus, terminal and medium
CN121501293B