An automatic positioning method and system for application program abnormal crash
Patent Information
- Application Number
- CN202611218419.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-12
- Publication Date
- 2026-09-29
AI Technical Summary
具体而言,当崩溃恰好发生在解释器所加载的脚本逻辑与原生编译函数相互调用的边界位置,且相关原生模块刚经由云端下发的热补丁被改写时,崩溃瞬间捕获的调用记录中会同时混杂依赖不同版本符号表的两类执行帧,常规的符号解析服务仅能基于设备当前上报的主程序版本选择单一的符号文件,无法识别并区分同一个调用链中因热更新产生的版本差异分支,导致崩溃报告中将不同语义层的帧错误地混合展现,最终输出难以对应真实业务逻辑的混淆结果,使后续的缺陷定位失去准确性
[0050]1.本发明通过将调用记录中的指令地址按解释器引擎版本标识和原生热补丁版本标识分别归集为两个地址集合,利用核密度估计捕捉两类地址在内存空间中的分布重叠程度,使得原本仅凭版本号逐帧比对难以发现的混合调用段能够被自动识别出来,无需依赖人工判断或预先配置的版本兼容列表,定位触发条件更贴合热更新引发的真实混叠特征。
Smart Images

Figure CN122838286A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer software error localization technology, and more specifically, to an automatic localization method and system for application crashes. Background Technology
[0002] In cloud-edge collaborative application architectures, applications often employ a hybrid approach, combining interpreted scripts with native system code. This allows for dynamic updates to business logic via cloud-based script distribution or hot patches. Such updates eliminate the need for redeploying the entire application, enabling rapid responses to evolving functional requirements at the edge. When a program crashes, the operating system or runtime environment generates a crash log containing information such as instruction addresses and function call sequences. The cloud collects these logs and maps the addresses to readable function names and locations using a symbol table, thus assisting developers in locating defects.
[0003] In existing crash information collection and symbol resolution processes, there are semantic mismatch defects caused by the interaction between hybrid runtime architecture and hot update mechanisms. Specifically, when a crash occurs precisely at the boundary between the script logic loaded by the interpreter and the native compiled functions, and the relevant native modules have just been rewritten via a hot patch distributed from the cloud, the call records captured at the moment of the crash will simultaneously contain two types of execution frames that depend on different versions of symbol tables. Conventional symbol resolution services can only select a single symbol file based on the main program version currently reported by the device, and cannot identify and distinguish the version difference branches in the same call chain caused by hot updates. This results in the crash report incorrectly mixing frames from different semantic layers, ultimately producing a confusing result that is difficult to correspond to the actual business logic, making subsequent defect localization inaccurate. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides an automatic location method and system for application abnormal crashes to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] An automatic method for locating application crashes includes:
[0007] S1: Obtain the call record uploaded when the application crashes. The call record contains the instruction address sequence and the interpreter engine version identifier and native hot patch version identifier corresponding to each instruction address.
[0008] S2: Assign instruction addresses associated with interpreter engine version identifiers to the first address set, and assign instruction addresses associated with native hot patch version identifiers to the second address set. Perform kernel density estimation on the first address set and the second address set respectively to obtain the first density distribution and the second density distribution. Calculate the distribution overlap between the first density distribution and the second density distribution. If the distribution overlap exceeds a preset threshold, it is determined that there is a mixed call segment.
[0009] S3: When there is a mixed call segment, extract the address encoding timestamp of the instruction address in the first address set, align and compare the address encoding timestamp with the hot patch activation window corresponding to the native hot patch version identifier, obtain the first symbol table from the cloud symbol service based on the comparison result, perform symbol resolution on the first address set, and obtain the interpreter frame symbol;
[0010] S4: Use the native hot patch version identifier to obtain the second symbol table from the cloud symbol service, perform symbol resolution on the second address set, and obtain the native frame symbols;
[0011] S5: Interpreter frame symbols and native frame symbols are alternately concatenated according to the instruction address sequence to form a complete call chain with unified version semantics in order to locate the crash point.
[0012] Furthermore, S1 includes:
[0013] After intercepting the application's abnormal exit signal, the value of the current instruction pointer register is read from the process's address space as the crash start address;
[0014] Starting from the crash start address, the call stack is traversed by backtracking through the stack frames, and the return addresses of each stack frame are extracted in turn to form an instruction address sequence;
[0015] Traverse each instruction address in the instruction address sequence, query the load attribute of the memory region to which the instruction address belongs, and if the load attribute is marked as executable and the file to which it belongs is the interpreter engine file, then associate the instruction address with the interpreter engine version identifier;
[0016] If the load attribute is marked as executable and the file it belongs to is a native shared library modified by a hot patch, then associate the address of the instruction with the native hot patch version identifier;
[0017] The instruction address sequence, the interpreter engine version identifier corresponding to each instruction address, and the native hot patch version identifier are combined into a call record and uploaded to the cloud.
[0018] Furthermore, S2 includes:
[0019] Iterate through the interpreter engine version identifier and native hot patch version identifier corresponding to each instruction address in the call record, add the instruction addresses with interpreter engine version identifiers to the first address set, and add the instruction addresses with native hot patch version identifiers to the second address set;
[0020] A Gaussian kernel function is selected, and kernel density estimation is performed on each instruction address in the first address set to obtain a first density distribution. Kernel density estimation is performed on each instruction address in the second address set to obtain a second density distribution.
[0021] Calculate the Barcol distance between the first density distribution and the second density distribution, and use the Barcol distance as the distribution overlap.
[0022] If the Bach distance is less than a preset threshold, it is determined that the instruction addresses of the first address set and the second address set are aliased in the address space, and it is confirmed that there is a mixed call segment in the call record.
[0023] Furthermore, selecting the Gaussian kernel function includes: obtaining the numerical distribution range of instruction addresses in the first address set, dividing the numerical distribution range into equal-width intervals, counting the frequency of instruction addresses in each equal-width interval, and calculating the standard deviation of the frequency as the distribution dispersion of the first address set; using the distribution dispersion as the bandwidth parameter of the Gaussian kernel function, the Gaussian kernel function is expressed as an exponential decay function related to the instruction address difference, and the bandwidth parameter controls the exponential decay rate.
[0024] Furthermore, S3 includes:
[0025] When it is determined that a mixed call segment exists, the address encoding timestamps embedded in the compilation stage of each instruction address in the first address set are read. The address encoding timestamps record the generation time of the code segment corresponding to the instruction address.
[0026] Get the hot patch activation window corresponding to the native hot patch version identifier. The hot patch activation window contains a time interval consisting of the hot patch distribution time and the hot patch loading completion time.
[0027] Each address encoding timestamp is aligned and compared with the hot patch activation window. If the address encoding timestamp is earlier than the hot patch issuance time, the corresponding instruction address is determined to be the address of the version before the hot patch, and the first symbol table corresponding to the version before the hot patch is obtained from the cloud symbol service.
[0028] If the address encoding timestamp is later than the time when the hot patch is loaded, the corresponding instruction address is determined to be the address of the version after the hot patch, and the first symbol table corresponding to the version after the hot patch is obtained from the cloud symbol service;
[0029] The first symbol table is used to perform symbol resolution on each instruction address in the first address set to obtain the interpreter frame symbol.
[0030] Furthermore, the address-encoded timestamp is generated as follows: when compiling and generating the code segment corresponding to the instruction address, the epoch timestamp of the compilation time is written into the reserved timestamp record position in the code segment. The timestamp record position has a fixed offset from the instruction address. When reading the instruction address in the first address set, the timestamp reading address is calculated through the fixed offset between the instruction address and the timestamp record position, and the address-encoded timestamp is extracted from the timestamp reading address.
[0031] Furthermore, S4 includes:
[0032] The native hot patch version identifier is used as the query key to send a symbol table request to the cloud symbol service. The cloud symbol service finds a second symbol table that strictly matches the native hot patch version identifier and returns it.
[0033] Receive the second symbol table returned by the cloud symbol service. The second symbol table contains the mapping relationship between instruction addresses and function names and line numbers.
[0034] Iterate through each instruction address in the second address set, look up the function name and line number corresponding to the instruction address in the second symbol table, and use the found function name and line number as the native frame symbol;
[0035] If any instruction address in the second address set fails to be found in the second symbol table, the instruction address is marked as an unresolved address, and the instruction address value of the unresolved address is used as a native frame symbol.
[0036] Furthermore, S5 includes:
[0037] Read the instruction address sequence from the call log. The instruction address sequence records the original order of the return addresses of each stack frame when the crash occurred.
[0038] Traverse each instruction address in the instruction address sequence. If the instruction address belongs to the first address set, then take the corresponding symbol from the interpreter frame symbol in order and fill it into the current position of the call chain.
[0039] If the instruction address belongs to the second address set, then the corresponding symbol is taken out from the native frame symbols in order and filled into the current position of the call chain;
[0040] All the symbols filled in sequence are concatenated to form a complete call chain with unified version semantics. The symbols in each frame of the complete call chain are arranged in the actual call order when the crash occurs.
[0041] The location of the first frame symbol in the complete call chain that exhibits abnormal semantics is identified as the crash point.
[0042] Furthermore, determining the location of the first frame symbol in the complete call chain that exhibits abnormal semantics as the crash point includes: traversing the function names of each frame symbol in the complete call chain; if the function name of the frame symbol contains the name of an exception signal handling function or a termination call function, then the frame symbol is determined to have abnormal semantics; if the function name of the frame symbol contains the name of a memory access function and the line number points to a load or store instruction, then the frame symbol is determined to have abnormal semantics; and determining the sequence number of the first frame symbol exhibiting abnormal semantics in the complete call chain as the crash point.
[0043] On the other hand, the present invention provides an automatic location system for application crashes, comprising:
[0044] The call log module is used to obtain the call log uploaded when the application crashes. The call log contains the instruction address sequence and the interpreter engine version identifier and native hot patch version identifier corresponding to each instruction address.
[0045] The mixed segment determination module is used to classify the instruction addresses associated with the interpreter engine version identifier into the first address set, classify the instruction addresses associated with the native hot patch version identifier into the second address set, perform kernel density estimation on the first address set and the second address set respectively to obtain the first density distribution and the second density distribution, calculate the distribution overlap between the first density distribution and the second density distribution, and if the distribution overlap exceeds a preset threshold, it is determined that there is a mixed call segment.
[0046] The alignment and resolution module is used to extract the address encoding timestamp of the instruction address in the first address set when there is a mixed call segment, align and compare the address encoding timestamp with the hot patch activation window corresponding to the native hot patch version identifier, obtain the first symbol table from the cloud symbol service based on the comparison result, and perform symbol resolution on the first address set to obtain the interpreter frame symbol.
[0047] The symbol resolution module is used to obtain the second symbol table from the cloud symbol service using the native hot patch version identifier, and to perform symbol resolution on the second address set to obtain the native frame symbols;
[0048] The call chain concatenation module is used to alternately concatenate interpreter frame symbols and native frame symbols according to the instruction address sequence to form a complete call chain with unified version semantics in order to locate the crash point.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] 1. This invention categorizes instruction addresses in the call record into two address sets based on the interpreter engine version identifier and the native hot patch version identifier. It then uses kernel density estimation to capture the degree of overlap in the distribution of these two types of addresses in memory space. This allows mixed call segments that were previously difficult to detect by comparing version numbers frame by frame to be automatically identified. It eliminates the need for manual judgment or a pre-configured version compatibility list, and the location triggering conditions are more consistent with the real aliasing characteristics caused by hot updates.
[0051] 2. After identifying the mixed call segment, the address encoding timestamp embedded in the instruction address is further extracted. The address encoding timestamp is aligned and compared with the hot patch effective window. According to the comparison result, the first symbol table of the corresponding time interval is selected to parse the interpreter frame. At the same time, the native hot patch version identifier is used to obtain the strictly matching second symbol table to parse the native frame. Then, the parsed interpreter frame symbols and native frame symbols are alternately concatenated according to the original order of the instruction address to restore the complete call chain with unified version semantics. This avoids semantic confusion caused by the cross-use of different version symbol tables due to hot patch version switching. It makes the crash location result directly correspond to the real business code before and after the patch, improving the accuracy of crash analysis of hybrid runtime applications under the cloud-edge collaborative architecture. Attached Figure Description
[0052] Figure 1 This is a flowchart of an automatic method for locating application crashes according to the present invention;
[0053] Figure 2 This is a schematic diagram of the structure of an automatic location system for abnormal application crashes according to the present invention. Detailed Implementation
[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0055] Example 1: Figure 1 This invention provides an automatic method for locating application crashes, comprising:
[0056] S1: Obtain the call record uploaded when the application crashes. The call record contains the instruction address sequence and the interpreter engine version identifier and native hot patch version identifier corresponding to each instruction address.
[0057] S2: Assign instruction addresses associated with interpreter engine version identifiers to the first address set, and assign instruction addresses associated with native hot patch version identifiers to the second address set. Perform kernel density estimation on the first address set and the second address set respectively to obtain the first density distribution and the second density distribution. Calculate the distribution overlap between the first density distribution and the second density distribution. If the distribution overlap exceeds a preset threshold, it is determined that there is a mixed call segment.
[0058] S3: When there is a mixed call segment, extract the address encoding timestamp of the instruction address in the first address set, align and compare the address encoding timestamp with the hot patch activation window corresponding to the native hot patch version identifier, obtain the first symbol table from the cloud symbol service based on the comparison result, perform symbol resolution on the first address set, and obtain the interpreter frame symbol;
[0059] S4: Use the native hot patch version identifier to obtain the second symbol table from the cloud symbol service, perform symbol resolution on the second address set, and obtain the native frame symbols;
[0060] S5: Interpreter frame symbols and native frame symbols are alternately concatenated according to the instruction address sequence to form a complete call chain with unified version semantics in order to locate the crash point.
[0061] In the implementation of S1, when an application crashes abnormally, the operating system kernel sends an abnormal exit signal to the application process. The pre-registered signal handler within the application process intercepts the abnormal exit signal, suspends the execution of other threads within the process, and reads the value of the current instruction pointer register from the process address space. The address stored in the instruction pointer register is the location of the instruction that triggered the abnormal exit signal; this address is used as the crash start address. Starting from the crash start address, the call stack is traversed by backtracking through stack frames. The stack frame backtracking process involves reading the value of the frame pointer register stored in the stack frame corresponding to the crash start address, obtaining the return address of the previous stack frame based on the base address pointed to by the frame pointer register, adding the return address to the instruction address sequence, and continuing to trace upwards using the frame pointer register value of the previous stack frame until all stack frames in the call stack have been traversed, forming the instruction address sequence.
[0062] When traversing each instruction address in the instruction address sequence, the load attributes of the memory region to which the instruction address belongs are queried. The load attributes of the memory region are written to the process's virtual memory region structure by the operating system kernel when loading executable files and shared libraries. The virtual memory region structure records the start address, end address, and permission flags of the memory region. It also records the file mapping information corresponding to the memory region, including the file path and file inode number. Searching for the instruction address within the virtual memory region structure, if the instruction address falls within the start and end address range of a certain virtual memory region structure, the permission flags and file mapping information of the virtual memory region structure are read. If the permission flag indicates executable and the file path in the file mapping information points to the interpreter engine file, the instruction address is associated with the interpreter engine version identifier. The interpreter engine version identifier is extracted from the metadata segment of the interpreter engine file. During compilation, the interpreter engine file writes a version number string to the reserved version identifier record location in the metadata segment. The interpreter engine file pointed to by the file path is read, and the interpreter engine version identifier is extracted from the version identifier record location. If the permission flag indicates executable, the file path in the file mapping information points to the native shared library, and the native shared library's file metadata contains a hot patch modification flag, then the instruction address is associated with the native hot patch version identifier. The hot patch modification flag is written to the native shared library's metadata segment by the cloud-based patch management system when a hot patch file is downloaded. The presence of the hot patch modification flag indicates that the native shared library has been modified by a hot patch. The native hot patch version identifier is extracted from the metadata segment of the hot-patched native shared library. During compilation, the native shared library writes the version number string into the reserved version identifier record location in the metadata segment. When a hot patch is applied, the version number string in the version identifier record location is updated to reflect the modified version.
[0063] After associating each instruction address in the instruction address sequence with its corresponding interpreter engine version identifier and native hot patch version identifier, a call record is formed. A call record is a structured data collection containing a list of instruction address sequences, a list of interpreter engine version identifiers corresponding to each instruction address in the sequence, and a list of native hot patch version identifiers corresponding to each instruction address. When an instruction address is not associated with an interpreter engine version identifier, the corresponding position is filled with an empty value; when an instruction address is not associated with a native hot patch version identifier, the corresponding position is filled with an empty value. After combination, the call record is uploaded to the cloud via a data transmission channel between the edge device and the cloud. This data transmission channel is a pre-established encrypted network connection between the edge device and the cloud. Before uploading, the call record undergoes serialization, converting the structured data collection into a byte stream before sending it to the cloud via the encrypted network connection. The cloud receiving end receives the byte stream, deserializes it, and recovers the call record for subsequent steps.
[0064] In the specific implementation of S2, the interpreter engine version identifier and native hot patch version identifier corresponding to each instruction address in the call record are traversed. Instruction addresses with an interpreter engine version identifier are added to the first address set, and instruction addresses with a native hot patch version identifier are added to the second address set. This process reads the interpreter engine version identifier list and the native hot patch version identifier list from the call record recovered from cloud deserialization. The interpreter engine version identifier list and the native hot patch version identifier list correspond one-to-one with each instruction address in the instruction address sequence list. The interpreter engine version identifier corresponding to each instruction address is checked sequentially. If the interpreter engine version identifier is not null, the instruction address is added to the first address set. The native hot patch version identifier corresponding to each instruction address is checked sequentially. If the native hot patch version identifier is not null, the instruction address is added to the second address set. An instruction address may be associated with both an interpreter engine version identifier and a native hot patch version identifier. In this case, the instruction address is added to both the first and second address sets simultaneously.
[0065] A Gaussian kernel function is selected. Kernel density estimation is performed on each instruction address in the first address set to obtain a first density distribution. Similarly, kernel density estimation is performed on each instruction address in the second address set to obtain a second density distribution. The Gaussian kernel function is a function that takes instruction address as its independent variable and outputs kernel density values. The value of the Gaussian kernel function is determined by the difference between the instruction address and the sample instruction address in the set, as well as the bandwidth parameter. The bandwidth parameter controls the smoothness of the Gaussian kernel function. A smaller bandwidth parameter results in faster weight decay around the sample instruction address, while a larger bandwidth parameter results in slower weight decay. The bandwidth parameter is obtained by acquiring the numerical distribution range of instruction addresses in the first address set, dividing the numerical distribution range into equal-width intervals, counting the frequency of instruction addresses within each equal-width interval, calculating the standard deviation of the frequencies as the distribution dispersion of the first address set, and using the distribution dispersion as the bandwidth parameter of the Gaussian kernel function. For example, the numerical distribution range of instruction addresses in the first address set is from 0x7f00 to 0x7fff in hexadecimal representation. This numerical distribution range is divided into multiple intervals of equal width, each interval being 0x20 in hexadecimal representation. The number of instruction addresses falling into each interval is counted to obtain a set of frequency values. The standard deviation of the frequency values is calculated. If the standard deviation is, for example, 2.5, then 2.5 is used as the bandwidth parameter of the Gaussian kernel function. When performing kernel density estimation on the first address set, for each address point in the address space, with each sample instruction address in the first address set as the center, a Gaussian kernel function with a bandwidth parameter of, for example, 2.5 is applied to calculate the kernel density contribution value of each sample instruction address to the address point. The kernel density value of the address point is obtained by summing the kernel density contribution values of all sample instruction addresses to the address point. After traversing all address points in the address space, the first density distribution is obtained. When performing kernel density estimation on the second address set, the same address space range and address point division method as the first address set are used. Taking each sample instruction address in the second address set as the center, the bandwidth parameter is also applied to calculate the kernel density contribution value of each sample instruction address to the address point and sum them to obtain the second density distribution.
[0066] The Barthold distance between the first and second density distributions is calculated and used as the distribution overlap. Barthold distance is a measure of the similarity between two probability distributions. A Barthold distance value ranges from 0 to 1; a distance closer to 0 indicates greater similarity, while a distance closer to 1 indicates greater difference. The Barthold distance is calculated by multiplying the square root of the kernel density value of the first density distribution at each address point by the square root of the kernel density value of the second density distribution at the corresponding address point. The sum of these multiplications over the address space yields the Barthold coefficient. Taking the natural logarithm of the coefficient and then taking its negative value gives the Barthold distance. As the distribution overlap is measured, a smaller Barthold distance indicates a higher degree of overlap between the first and second density distributions in the address space; that is, a more severe aliasing exists between the instruction addresses associated with the interpreter engine version identifier and the instruction addresses associated with the native hot patch version identifier.
[0067] If the Barcol distance is less than a preset threshold, it is determined that the instruction addresses of the first address set and the second address set are aliased in the address space, confirming the existence of mixed call segments in the call records. The preset threshold is determined based on the distribution characteristics of the instruction addresses associated with the interpreter engine version identifier and the instruction addresses associated with the native hot patch version identifier in the address space under normal operating conditions. The preset threshold is set by collecting call records of multiple crash samples under normal operating conditions where the application has not undergone hot patch updates. From the call records of each crash sample, the instruction addresses associated with the interpreter engine version identifier and the instruction addresses associated with the native hot patch version identifier are extracted. The Barcol distance corresponding to each call record is calculated according to the method in this step, resulting in a set of normal state Barcol distance values. Under normal conditions, the instruction addresses associated with the interpreter engine version identifier and the instruction addresses associated with the native hot patch version identifier are usually distributed in different non-overlapping regions in the address space, and the normal state Barcol distance value tends to be close to 1, for example, the average value of the normal state Barcol distance value is 0.95. Based on this, a preset threshold is set. The preset threshold is the boundary value that distinguishes the degree of overlap between the normal state distribution and the degree of overlap between the aliased state distribution. The preset threshold is set to, for example, 0.7. When the Bach distance calculated from a call record is less than a preset threshold, such as 0.7, it indicates that the instruction addresses associated with the interpreter engine version identifier and the instruction addresses associated with the native hot patch version identifier have an unnatural overlap and intersection in the address space. In other words, the instruction addresses of the first address set and the second address set are aliased in the address space, confirming the existence of a mixed call segment in the call record. The confirmation of the mixed call segment indicates that the instruction address sequence in the call record contains both interpreter frames and native frames, and that these frames are mixed together due to version differences caused by hot updates. Independent symbol resolution for each version is required to accurately reconstruct the crash scenario.
[0068] In the specific implementation of S3, when S2 determines that a mixed call segment exists, the address-encoded timestamps embedded during the compilation phase for each instruction address in the first address set are read. The address-encoded timestamps record the generation time of the code segment corresponding to the instruction address. The address-encoded timestamps are generated by writing the epoch timestamp of the compilation time into a reserved timestamp record position in the code segment when the code segment corresponding to the instruction address is generated during compilation. The timestamp record position has a fixed offset from the instruction address. This fixed offset is determined by the linker script during the compilation and linking phase. The linker script reserves a timestamp record slot in the code segment for each function entry instruction address, and the offset of the timestamp record slot relative to the function entry instruction address is a fixed offset. For example, the fixed offset is set to 8 bytes forward from the function entry instruction address in the code segment. The epoch timestamp of the compilation time is written into the timestamp record position as a 64-bit unsigned integer, representing the number of seconds elapsed since 00:00:00 UTC on January 1, 1970. When reading instruction addresses from the first address set, the timestamp read address is calculated using a fixed offset between the instruction address and the timestamp record position. The timestamp read address equals the instruction address plus the fixed offset. A 64-bit unsigned integer is extracted from the timestamp read address as the address-encoded timestamp. Each instruction address in the first address set is iterated over, and the timestamp read address is obtained by adding the fixed offset to the instruction address. The address-encoded timestamp is then read from the timestamp read address, and the read address-encoded timestamps are stored in correspondence with the instruction addresses, forming an address-encoded timestamp mapping for each instruction address in the first address set.
[0069] Retrieve the hot patch activation window corresponding to the native hot patch version identifier. The hot patch activation window comprises a time interval consisting of the hot patch distribution time and the hot patch loading completion time. The hot patch distribution time is the moment when the cloud-based patch management system pushes the hot patch file to the edge device. The cloud-based patch management system records the push time as the hot patch distribution time when pushing the hot patch file and sends it to the edge device along with the hot patch file's metadata. The hot patch loading completion time is the moment when the edge device's runtime environment loads the modified code in the hot patch file into the native shared library and completes the mapping in the process address space. The edge device's runtime environment records the loading completion time as the hot patch loading completion time after the hot patch is loaded. The correspondence between the native hot patch version identifier and the hot patch activation window is maintained by the cloud-based patch management system. The cloud-based patch management system assigns a native hot patch version identifier to the hot patch file each time a hot patch file is distributed, and associates and stores the native hot patch version identifier, the hot patch distribution time, and the hot patch loading completion time reported by the edge device as the hot patch activation window. When edge devices upload call records, they carry the native hot patch version identifier. The cloud queries the corresponding hot patch activation window from the patch management system based on the native hot patch version identifier.
[0070] Each address encoding timestamp is aligned and compared with the hot patch activation window. The alignment process involves comparing the address encoding timestamp with the hot patch release time within the hot patch activation window, and then comparing it with the hot patch loading completion time within the hot patch activation window. If the address encoding timestamp is earlier than the hot patch release time, the corresponding instruction address is determined to be the address of the version before the hot patch, and the first symbol table corresponding to the version before the hot patch is obtained from the cloud symbol service. The method for obtaining the first symbol table corresponding to the version before the hot patch is as follows: a symbol table request is sent to the cloud symbol service using the native hot patch version identifier and the previous hot patch version tag as a join query key. The cloud symbol service locates the symbol table set of the native shared library based on the native hot patch version identifier, searches for the symbol table generated most recently before the hot patch release time that has not been modified by the hot patch, and returns it as the first symbol table corresponding to the version before the hot patch. If the address encoding timestamp is later than the hot patch loading completion time, the corresponding instruction address is determined to be the address of the version after the hot patch, and the first symbol table corresponding to the version after the hot patch is obtained from the cloud symbol service. The method for obtaining the first symbol table corresponding to the hot-patched version is as follows: A symbol table request is sent to the cloud symbol service using the native hot-patched version identifier and the hot-patched version marker as a join query key. The cloud symbol service locates the symbol table set of the native shared library based on the native hot-patched version identifier. It then searches for the symbol table generated after the hot-patched loading completion time and containing the hot-patched modifications, which is taken as the first symbol table corresponding to the hot-patched version and returned. If the address encoding timestamp is between the hot-patched release time and the hot-patched loading completion time, the corresponding instruction address is determined to be the hot-patched execution version address. The hot-patched execution version address indicates that the code segment corresponding to the instruction address was generated within the time window after the hot-patched release but before it has finished loading. The most recently compiled symbol table generated before the hot-patched release time is obtained from the cloud symbol service as the first symbol table.
[0071] The first symbol table is used to perform symbol resolution on each instruction address in the first address set to obtain the interpreter frame symbol. The symbol resolution process involves using the instruction address as the lookup key to search the first symbol table, which contains a mapping between instruction addresses and function names and line numbers. If a matching entry exists in the first symbol table, the function name and line number from the matching entry are extracted as the interpreter frame symbol, which includes a function name field and a line number field. Different instruction addresses in the first address set may correspond to the pre-hot-patched version address, the post-hot-patched version address, or the version address during hot-patched execution. Symbol resolution for these different versions of instruction addresses is performed using the corresponding version of the first symbol table. After resolving all instruction addresses in the first address set, the interpreter frame symbol set is obtained, and each interpreter frame symbol in the interpreter frame symbol set maintains a one-to-one correspondence with each instruction address in the first address set.
[0072] In the implementation of S4, the native hot patch version identifier is used as the query key to send a symbol table request to the cloud symbol service. The cloud symbol service searches for and returns a second symbol table that strictly matches the native hot patch version identifier. The native hot patch version identifier is extracted from the metadata segment of the native shared library modified by the hot patch in step S1 and uploaded to the cloud along with the instruction address sequence in the call record. After receiving the symbol table request, the cloud symbol service parses the native hot patch version identifier carried in the request and searches for it in the symbol table index maintained by the cloud symbol service. The symbol table index maintained by the cloud symbol service is a key-value mapping structure with the native hot patch version identifier as the key and the symbol table storage path as the value. Each native hot patch version identifier in the symbol table index uniquely corresponds to a symbol table storage path, which points to the second symbol table file stored on the cloud storage device. The cloud symbol service generates the second symbol table file when compiling the native shared library and writes it into the symbol table index after associating the native hot patch version identifier with the storage path of the second symbol table file. Once the hot patch on the edge device is loaded, the cloud-based patch management system updates the mapping between the native hot patch version identifier and the hot patch activation window. Simultaneously, it synchronizes the updated native hot patch version identifier to the cloud symbol service, which marks the second symbol table file corresponding to that version as the current valid version. The cloud symbol service locates the matching symbol table storage path in the symbol table index based on the native hot patch version identifier, reads the second symbol table file from the storage path, and returns its contents as a response to the requester. The second symbol table file is a text file containing multiple lines of records. Each line contains the instruction address, the function name corresponding to the instruction address, and the line number corresponding to the instruction address, separated by a delimiter.
[0073] The system receives a second symbol table returned by a cloud-based symbol service. This second symbol table contains mappings from instruction addresses to function names and line numbers. The second symbol table returned by the cloud symbol service is serialized text content. This serialized text content is parsed into mappings from instruction addresses to function names and line numbers. The parsing process involves splitting the text content of the second symbol table line by line. Each line of text is then split into instruction address, function name, and line number fields according to a delimiter. These fields are stored as mapping entries, and all mapping entries constitute the mapping relationship between instruction addresses and function names and line numbers. The mapping relationship is stored in memory as a lookup table, using the instruction address as the lookup key. This table supports fast retrieval of the corresponding function name and line number based on the instruction address. The lookup table is stored as an array sorted by instruction address. Each element in the array contains three members: instruction address, function name, and line number. A binary search method is used to locate the element corresponding to the instruction address in the array.
[0074] The process iterates through each instruction address in the second address set, searching for the corresponding function name and line number in the second symbol table. The found function name and line number are then used as the native frame symbol. The second address set is the set of addresses into which instruction addresses with native hot patch version identifiers are grouped by traversing the call records in step S2. During the traversal of the second address set, each instruction address is sequentially retrieved from the set, and a binary search is performed in the lookup table of the second symbol table using the retrieved instruction address as the lookup key. The binary search process involves comparing the instruction address with the instruction address of the middle element in the lookup table. If they are equal, the search is successful; if the instruction address is less than the middle element's instruction address, the search continues in the first half of the lookup table; if the instruction address is greater than the middle element's instruction address, the search continues in the second half of the lookup table. Upon successful search, the function name and line number are extracted from the corresponding mapping entry and combined to form the native frame symbol. The native frame symbol uses the same data structure as the interpreter frame symbol, containing a function name field and a line number field. The function name field stores the function name string, and the line number field stores the line number value. The found native frame symbols are stored in the native frame symbol sequence in the order of the instruction addresses in the second address set. Each native frame symbol in the native frame symbol sequence has a one-to-one correspondence with each instruction address in the second address set.
[0075] If any instruction address in the second address set fails to find a match in the second symbol table, the instruction address is marked as an unresolved address, and its address value is used as the native frame symbol. A search failure occurs when the instruction address has no matching entry in the lookup table. If a binary search fails to find an entry equal to the instruction address in the lookup table, the search is considered unsuccessful. In this case, the instruction address is marked as an unresolved address, and a corresponding native frame symbol is generated. The function name field of the native frame symbol is filled with a preset identifier indicating unresolvedness. This preset identifier is a special string, such as a string composed of a preset prefix and the hexadecimal representation of the instruction address. The line number field of the native frame symbol is filled with a preset value, such as 0. The purpose of marking the instruction address as an unresolved address and using its address value as the native frame symbol is to preserve the position information of that frame in the crash call chain, making it easier for developers to identify frames with missing symbols and take subsequent measures such as supplementing the symbol table or manual analysis. After the unresolved address is processed, the generated native frame symbols are stored in the native frame symbol sequence according to the order of the instruction address in the second address set. Together with the native frame symbols that were successfully found, they form a complete native frame symbol sequence, which is used by the subsequent S5 steps.
[0076] In the implementation of S5, the instruction address sequence is read from the call record. This sequence records the original order of the return addresses of each stack frame at the time of the crash. The call record is uploaded to the cloud by the edge device in step S1 and deserialized and restored before step S2. The instruction address sequence is an ordered list contained in the call record, where each element is an instruction address. The order of these addresses represents the call order of each stack frame from the innermost to the outermost layer at the time of the crash. Reading the instruction address sequence involves extracting the instruction address sequence list from the call record's data structure. This list is stored as an array, where smaller index positions correspond to stack frames closer to the crash trigger point, and larger index positions correspond to stack frames further away from the crash trigger point. The original order of the instruction address sequence is naturally formed during the backtracking of the call stack in step S1. The backtracking process starts from the crash start address and proceeds upwards; therefore, the first instruction address in the sequence is the crash start address, and subsequent instruction addresses are the return addresses of each call stack layer.
[0077] Traverse each instruction address in the instruction address sequence. If the instruction address belongs to the first address set, then retrieve the corresponding symbol from the interpreter frame symbols in order and fill it into the current position of the call chain. The method for determining whether an instruction address belongs to the first address set is to match the instruction address with the instruction addresses in the first address set. The first address set was completed in step S2 and stores all instruction addresses associated with the interpreter engine version identifier. Since the order of instruction addresses in the first address set is consistent with the order of occurrence of instruction addresses belonging to the first address set in the instruction address sequence, the order of interpreter frame symbols in the interpreter frame symbol sequence also corresponds one-to-one with the order of instruction addresses in the first address set. The interpreter frame symbol sequence is generated in step S3, and each interpreter frame symbol in the interpreter frame symbol sequence maintains the same order as each instruction address in the first address set. The specific method for retrieving corresponding symbols sequentially from the interpreter frame symbols is as follows: A pointer is maintained pointing to the current position of the interpreter frame symbol sequence. Initially, the pointer points to the first interpreter frame symbol in the sequence. When an instruction address in the instruction address sequence belongs to the first address set, the interpreter frame symbol currently pointed to by the pointer is read and filled into the current vacant position in the call chain. After filling, the pointer is moved one position forward to point to the next interpreter frame symbol, for use by subsequent instruction addresses belonging to the first address set. Filling into the current position of the call chain means placing the interpreter frame symbol at the same index position in the call chain, ensuring that the order of frame symbols in the call chain is completely consistent with the order of the instruction address sequence. The interpreter frame symbol contains a function name field and a line number field. During filling, the values of the function name field and the line number field are written verbatim to the corresponding positions in the call chain.
[0078] If the instruction address belongs to the second address set, the corresponding symbol is retrieved from the native frame symbols in sequence and filled into the current position of the call chain. The method for determining if an instruction address belongs to the second address set is to match the instruction address with the instruction addresses in the second address set. The second address set was completed in step S2 and stores all instruction addresses associated with native hot patch version identifiers. The order of native frame symbols in the native frame symbol sequence corresponds one-to-one with the order of instruction addresses in the second address set. The native frame symbol sequence is generated in step S4. The specific method for retrieving the corresponding symbol from the native frame symbols in sequence is as follows: a pointer is maintained pointing to the current position of the native frame symbol sequence. The pointer initially points to the first native frame symbol in the sequence. When an instruction address in the instruction address sequence belongs to the second address set, the native frame symbol currently pointed to by the pointer is read, and this symbol is filled into the current empty position of the call chain. After filling, the pointer is moved one position forward to point to the next native frame symbol. The native frame symbol also contains a function name field and a line number field. When filling into the call chain, the values of the function name field and the line number field are written as is. When an instruction address belongs to both the first address set and the second address set, the instruction address has been added to both sets in step S2. At this time, according to the definition of the hybrid call segment, the instruction address has the characteristics of both the interpreter frame and the native frame. The processing method is to first take the symbol from the interpreter frame symbol sequence and fill it in, or to decide according to the configuration. In this embodiment, the interpreter frame symbol is filled in first, and the pointer of the native frame symbol sequence is synchronously shifted one position to the right to maintain consistency. This is because the core problem of the hybrid call segment is that the interpreter frame symbol is incorrectly mixed due to hot updates. Therefore, using the interpreter frame semantics as the basis is more conducive to restoring the unified version semantics.
[0079] All symbols entered in sequence are concatenated to form a complete call chain with unified version semantics. The symbols in each frame of the complete call chain are arranged according to the actual call order at the time of the crash. The concatenation operation appends the function name and line number fields at each position in the call chain to form a readable call frame string. The string format of each frame is function name followed by line number, separated by a colon, with frames separated by newline characters. The complete call chain formed by concatenation is a continuous text sequence. The first line of the text sequence corresponds to the crash trigger point, and the last line corresponds to the outermost function of the call stack. Each frame of the complete call chain contains correct symbols after version alignment, and there is no semantic confusion caused by differences in hot-update versions; therefore, this complete call chain has unified version semantics.
[0080] The location of the first frame symbol in the complete call chain exhibiting abnormal semantics is determined as the crash point. The function names of each frame symbol in the complete call chain are traversed. If the function name of a frame symbol contains an exception signal handling function name or a termination call function name, then the frame symbol is considered to have abnormal semantics. Exception signal handling function names include the entry names of predefined signal handling functions in the operating system, such as sigsegv_handler, sigbus_handler, sigill_handler, etc., as well as the names of default exception handling functions registered in the runtime library. Termination call function names include the names of functions called when the program terminates voluntarily, such as abort, exit, terminate, etc. If the function name of a frame symbol contains a memory access function name and the line number points to a load or store instruction, then the frame symbol is considered to have abnormal semantics. Memory access function names include string manipulation function names such as memcpy, strcpy, memmove, as well as memory allocation and deallocation function names such as malloc, free. The method for determining whether a line number points to a load or store instruction is as follows: Machine instruction bytes are read from the program storage medium based on the instruction address corresponding to the frame symbol. These machine instruction bytes are then disassembled and parsed to extract the opcode. If the opcode belongs to a load or store instruction, the line number points to either. Load instructions include those that load data from memory into registers, while store instructions include those that store data from registers into memory. If neither the function name nor the line number of the frame symbol meets the above conditions, the process continues to traverse the next frame symbol. After determining that an abnormal semantic has occurred, the sequence number of the first frame symbol with abnormal semantics in the complete call chain is determined as the crash point. The sequence number is counted starting from the first frame of the complete call chain, with the first frame number being 1, and incrementing sequentially. If no abnormal semantic has occurred after traversing all frame symbols, the first frame of the complete call chain is determined as the crash point.
[0081] Example 2: Figure 2 A schematic diagram of an automatic application crash location system according to the present invention is provided. The automatic application crash location system includes:
[0082] The call log module is used to obtain the call log uploaded when the application crashes. The call log contains the instruction address sequence and the interpreter engine version identifier and native hot patch version identifier corresponding to each instruction address.
[0083] The mixed segment determination module is used to classify the instruction addresses associated with the interpreter engine version identifier into the first address set, classify the instruction addresses associated with the native hot patch version identifier into the second address set, perform kernel density estimation on the first address set and the second address set respectively to obtain the first density distribution and the second density distribution, calculate the distribution overlap between the first density distribution and the second density distribution, and if the distribution overlap exceeds a preset threshold, it is determined that there is a mixed call segment.
[0084] The alignment and resolution module is used to extract the address encoding timestamp of the instruction address in the first address set when there is a mixed call segment, align and compare the address encoding timestamp with the hot patch activation window corresponding to the native hot patch version identifier, obtain the first symbol table from the cloud symbol service based on the comparison result, and perform symbol resolution on the first address set to obtain the interpreter frame symbol.
[0085] The symbol resolution module is used to obtain the second symbol table from the cloud symbol service using the native hot patch version identifier, and to perform symbol resolution on the second address set to obtain the native frame symbols;
[0086] The call chain concatenation module is used to alternately concatenate interpreter frame symbols and native frame symbols according to the instruction address sequence to form a complete call chain with unified version semantics in order to locate the crash point.
[0087] All calculations involved in the embodiments are performed using dimensionless numerical values, and the preset parameters and thresholds in the calculations can be set by those skilled in the art according to actual conditions.
[0088] This technical solution can be flexibly deployed, for example, as embedded software running on device hardware, or installed on personal computers or other smart terminals with user interfaces, thus adapting to various hardware environments and usage requirements.
[0089] The above solutions can be implemented in software, hardware, firmware, or a combination thereof. When implemented in software, they are presented, in whole or in part, as a computer program product, containing one or more computer instructions or programs. When these instructions or programs are loaded and executed on a computer, results are produced corresponding to the processes or functions of the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one medium to another via wired or wireless means, such as from a website, server, or data center via wired means such as fiber optic cables, twisted-pair cables, or coaxial cables, or wireless means such as infrared or microwaves to another site. A computer-readable storage medium refers to any usable medium that a computer can access or a data storage device such as a server or data center that contains one or more usable media, including magnetic media such as floppy disks, hard disks, and magnetic tapes, optical media such as DVDs, and semiconductor media such as solid-state drives.
[0090] The specific working process of the system, device and module can be found in the method embodiment, and will not be repeated here.
[0091] The disclosed systems, devices, and methods can be implemented in other ways. The device embodiments are for illustrative purposes only, and the module division is only a logical division. In practice, different divisions can be implemented, such as merging or integrating multiple modules or components, or omitting some features. Coupling, direct coupling, or communication connections between the components can be achieved through interfaces, while indirect coupling or communication connections can take electrical, mechanical, or other forms.
[0092] The modules described as separate components may or may not be physically separated. The components shown as modules can be physical hardware or software, and can be deployed centrally or distributed across multiple network nodes. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0093] The functional modules in each embodiment can be integrated into one processing module, or they can exist independently, or two or more modules can be integrated into one.
[0094] If the functionality is implemented as a software module and used as an independent product, it can be stored in a computer-readable storage medium. Under this understanding, the substantial contribution of the technical solution of this application can be embodied in the form of a software product. This computer software product is stored in a storage medium and contains instructions to cause a computer device, such as a personal computer, server, or network device, to execute all or part of the steps of the methods in the embodiments of this application. The storage medium includes any medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory, random access memory, magnetic disk, or optical disk.
[0095] The above are merely specific embodiments of this application, and the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should fall within the scope of protection of this application.
Claims
1. An automatic method for locating application crashes, characterized in that, include: S1: Obtain the call record uploaded when the application crashes. The call record contains the instruction address sequence and the interpreter engine version identifier and native hot patch version identifier corresponding to each instruction address. S2: Assign instruction addresses associated with interpreter engine version identifiers to the first address set, and assign instruction addresses associated with native hot patch version identifiers to the second address set. Perform kernel density estimation on the first address set and the second address set respectively to obtain the first density distribution and the second density distribution. Calculate the distribution overlap between the first density distribution and the second density distribution. If the distribution overlap exceeds a preset threshold, it is determined that there is a mixed call segment. S3: When there is a mixed call segment, extract the address encoding timestamp of the instruction address in the first address set, align and compare the address encoding timestamp with the hot patch activation window corresponding to the native hot patch version identifier, obtain the first symbol table from the cloud symbol service based on the comparison result, perform symbol resolution on the first address set, and obtain the interpreter frame symbol; S4: Use the native hot patch version identifier to obtain the second symbol table from the cloud symbol service, perform symbol resolution on the second address set, and obtain the native frame symbols; S5: Interpreter frame symbols and native frame symbols are alternately concatenated according to the instruction address sequence to form a complete call chain with unified version semantics in order to locate the crash point.
2. The automatic location method for application crashes according to claim 1, characterized in that, S1 includes: After intercepting the application's abnormal exit signal, the value of the current instruction pointer register is read from the process's address space as the crash start address; Starting from the crash start address, the call stack is traversed by backtracking through the stack frames, and the return addresses of each stack frame are extracted in turn to form an instruction address sequence; Traverse each instruction address in the instruction address sequence, query the load attribute of the memory region to which the instruction address belongs, and if the load attribute is marked as executable and the file to which it belongs is the interpreter engine file, then associate the instruction address with the interpreter engine version identifier; If the load attribute is marked as executable and the file it belongs to is a native shared library modified by a hot patch, then associate the address of the instruction with the native hot patch version identifier; The instruction address sequence, the interpreter engine version identifier corresponding to each instruction address, and the native hot patch version identifier are combined into a call record and uploaded to the cloud.
3. The automatic location method for application crashes according to claim 1, characterized in that, S2 include: Iterate through the interpreter engine version identifier and native hot patch version identifier corresponding to each instruction address in the call record, add the instruction addresses with interpreter engine version identifiers to the first address set, and add the instruction addresses with native hot patch version identifiers to the second address set; A Gaussian kernel function is selected, and kernel density estimation is performed on each instruction address in the first address set to obtain a first density distribution. Kernel density estimation is performed on each instruction address in the second address set to obtain a second density distribution. Calculate the Barcol distance between the first density distribution and the second density distribution, and use the Barcol distance as the distribution overlap. If the Bach distance is less than a preset threshold, it is determined that the instruction addresses of the first address set and the second address set are aliased in the address space, and it is confirmed that there is a mixed call segment in the call record.
4. The automatic location method for application crashes according to claim 3, characterized in that, The selection of the Gaussian kernel function includes: obtaining the numerical distribution range of instruction addresses in the first address set, dividing the numerical distribution range into equal-width intervals, counting the frequency of instruction addresses in each equal-width interval, and calculating the standard deviation of the frequency as the distribution dispersion of the first address set; using the distribution dispersion as the bandwidth parameter of the Gaussian kernel function, the Gaussian kernel function is expressed as an exponential decay function related to the instruction address difference, and the bandwidth parameter controls the exponential decay rate.
5. The automatic location method for application crashes according to claim 1, characterized in that, S3 include: When it is determined that a mixed call segment exists, the address encoding timestamps embedded in the compilation stage of each instruction address in the first address set are read. The address encoding timestamps record the generation time of the code segment corresponding to the instruction address. Get the hot patch activation window corresponding to the native hot patch version identifier. The hot patch activation window contains a time interval consisting of the hot patch distribution time and the hot patch loading completion time. Each address encoding timestamp is aligned and compared with the hot patch activation window. If the address encoding timestamp is earlier than the hot patch issuance time, the corresponding instruction address is determined to be the address of the version before the hot patch, and the first symbol table corresponding to the version before the hot patch is obtained from the cloud symbol service. If the address encoding timestamp is later than the time when the hot patch is loaded, the corresponding instruction address is determined to be the address of the version after the hot patch, and the first symbol table corresponding to the version after the hot patch is obtained from the cloud symbol service; The first symbol table is used to perform symbol resolution on each instruction address in the first address set to obtain the interpreter frame symbol.
6. The automatic location method for application crashes according to claim 5, characterized in that, The address-encoded timestamp is generated as follows: when compiling the code segment corresponding to the instruction address, the epoch timestamp of the compilation time is written into the reserved timestamp record position in the code segment. The timestamp record position has a fixed offset from the instruction address. When reading the instruction address in the first address set, the timestamp read address is calculated through the fixed offset between the instruction address and the timestamp record position, and the address-encoded timestamp is extracted from the timestamp read address.
7. The automatic location method for application crashes according to claim 1, characterized in that, S4 includes: The native hot patch version identifier is used as the query key to send a symbol table request to the cloud symbol service. The cloud symbol service finds a second symbol table that strictly matches the native hot patch version identifier and returns it. Receive the second symbol table returned by the cloud symbol service. The second symbol table contains the mapping relationship between instruction addresses and function names and line numbers. Iterate through each instruction address in the second address set, look up the function name and line number corresponding to the instruction address in the second symbol table, and use the found function name and line number as the native frame symbol; If any instruction address in the second address set fails to be found in the second symbol table, the instruction address is marked as an unresolved address, and the instruction address value of the unresolved address is used as a native frame symbol.
8. The automatic location method for application crashes according to claim 1, characterized in that, S5 include: Read the instruction address sequence from the call log. The instruction address sequence records the original order of the return addresses of each stack frame when the crash occurred. Traverse each instruction address in the instruction address sequence. If the instruction address belongs to the first address set, then take the corresponding symbol from the interpreter frame symbol in order and fill it into the current position of the call chain. If the instruction address belongs to the second address set, then the corresponding symbol is taken out from the native frame symbols in order and filled into the current position of the call chain; All the symbols filled in sequence are concatenated to form a complete call chain with unified version semantics. The symbols in each frame of the complete call chain are arranged in the actual call order when the crash occurs. The location of the first frame symbol in the complete call chain that exhibits abnormal semantics is identified as the crash point.
9. The automatic location method for application crashes according to claim 8, characterized in that, Determining the location of the first frame symbol in the complete call chain that exhibits abnormal semantics as the crash point involves: traversing the function names of each frame symbol in the complete call chain; if the function name of a frame symbol contains the name of an exception signal handling function or a termination call function, then the frame symbol is determined to have abnormal semantics; if the function name of a frame symbol contains the name of a memory access function and the line number points to a load or store instruction, then the frame symbol is determined to have abnormal semantics; and determining the sequence number of the first frame symbol exhibiting abnormal semantics in the complete call chain as the crash point.
10. An automatic system for locating application crashes, used to implement the automatic method for locating application crashes as described in any one of claims 1-9, characterized in that, include: The call log module is used to obtain the call log uploaded when the application crashes. The call log contains the instruction address sequence and the interpreter engine version identifier and native hot patch version identifier corresponding to each instruction address. The mixed segment determination module is used to classify the instruction addresses associated with the interpreter engine version identifier into the first address set, classify the instruction addresses associated with the native hot patch version identifier into the second address set, perform kernel density estimation on the first address set and the second address set respectively to obtain the first density distribution and the second density distribution, calculate the distribution overlap between the first density distribution and the second density distribution, and if the distribution overlap exceeds a preset threshold, it is determined that there is a mixed call segment. The alignment and resolution module is used to extract the address encoding timestamp of the instruction address in the first address set when there is a mixed call segment, align and compare the address encoding timestamp with the hot patch activation window corresponding to the native hot patch version identifier, obtain the first symbol table from the cloud symbol service based on the comparison result, and perform symbol resolution on the first address set to obtain the interpreter frame symbol. The symbol resolution module is used to obtain the second symbol table from the cloud symbol service using the native hot patch version identifier, and to perform symbol resolution on the second address set to obtain the native frame symbols; The call chain concatenation module is used to alternately concatenate interpreter frame symbols and native frame symbols according to the instruction address sequence to form a complete call chain with unified version semantics in order to locate the crash point.