Model-based updating of dial-back data
By employing a dynamic update method based on machine learning models, the problem of data collection gaps in the reference code mapping of callback data was solved, improving the accuracy and efficiency of defect analysis and enabling a more intelligent defect resolution process.
Patent Information
- Application Number
- CN202480071845.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2024-10-14
- Publication Date
- 2026-06-09
Smart Images

Figure CN122181127A_ABST
Abstract
Description
Background Technology
[0001] This disclosure relates to methods, apparatus, and products for model-based updating of call home data. Summary of the Invention
[0002] According to embodiments of this disclosure, various methods, apparatuses, and products for model-based updates of callback data are described herein. In some aspects, model-based updates of callback data include receiving problem analysis data associated with a computing system and identifying one or more reference codes within the problem analysis data that are associated with one or more defects. Based on the data associated with the one or more defects, a portion of the problem analysis data is determined to be used as training data for a model. A representation of each of the one or more reference codes within the portion of the problem analysis data is input into the model. The model is configured to output one or more data confidence scores based on the one or more reference codes. The association between the one or more reference codes and the one or more defects is updated based on the one or more data confidence scores. Attached Figure Description
[0003] Figure 1 An example computing environment based on various aspects of this disclosure is shown.
[0004] Figure 2 Another example computing environment according to various aspects of this disclosure is shown.
[0005] Figure 3 The present disclosure illustrates the use of various aspects thereof for Figure 2 An example of the processing flow in a computing environment.
[0006] Figure 4 A flowchart illustrating an example process for model-based updating of data rollback limits according to various aspects of this disclosure is shown.
[0007] Figure 5 A flowchart illustrating another example of a model-based update process for callback data, according to various aspects of this disclosure, is shown. Detailed Implementation
[0008] Problem analysis data is used to identify and correct defects (such as software or hardware errors) that may occur in relation to services or other programs provided by the computing system. Problem analysis data includes information related to the identified cause of the defect, such as files created or accessed during service provision, and is typically collected at the computing system by one or more agents during the execution of the service or program. For defects that cannot be easily fixed locally on the computing system, the problem analysis data is typically transferred to the service provider associated with the computing system for further analysis; this is known as “callback” communication. The service provider can then determine the steps required to correct the defect and communicate these steps to the customer associated with the computing system or perform them remotely to correct the defect.
[0009] In problem analysis, reference codes are callbacks mapped to various sets of data to aid in defect resolution. A reference code is a coded value indicating a diagnostic result of a computing system associated with a specific defect. For example, different types of hardware or software defects / bugs or diagnostic results can be assigned unique reference codes. The mapping from reference codes to data is typically determined by each reference code owner. However, sometimes data is only identified as necessary for defect resolution after the initial decision to take action on the problem has occurred. This leads to gaps in the data collection required to resolve the problem. Furthermore, with changes in or departure of reference code owners, the knowledge base and understanding of precisely identifying the data needed to resolve the defect may be lost.
[0010] Various embodiments of model-based updates for callback data in which errors or defects requiring action (e.g., callbacks) are identified, wherein problem analysis data collection regarding the error or defect is initiated. After a problem situation is identified as requiring action, data usage by the defect resolver is tracked to allow for the identification of data within the problem analysis data that is useful for resolving the error or defect. In a particular embodiment, heatmap tracking of data usage is used to reprioritize files or other data within the problem analysis data to identify data within the problem analysis data that is useful for resolving a specific defect. The identified problem analysis data is used as training data for the model. The model receives vector associations made from a long-term corpus of problem analysis data represented as reference codes as input and outputs a data confidence score. The output is dynamically fed back to the reference code-to-data mapping and pre-determined data, enabling future callbacks to utilize a more intelligently made set of problem analysis data. In one or more embodiments, the model is a machine learning model, such as a support vector machine, neural network, decision tree, or any other suitable machine learning model.
[0011] In a particular embodiment, the training model takes reference code vector associations as input; that is, each reference code is vectorized using features from training data tracked by heatmaps used in the data based on a question corpus, and outputs a data confidence score. In a particular embodiment, the features used for vectorizing the reference codes may include one or more fields from the question analysis data, related defect links (e.g., links created between related defects in the question analysis data), and / or sentence embeddings related to textual information found in the question analysis data. In one or more embodiments, the association between the reference codes and the data is dynamically updated, wherein the data confidence score is used to identify which sets of data in the question analysis data are essential for each reference code.
[0012] Figure 1 An example computing environment according to various aspects of this disclosure is illustrated. Computing environment 100 includes an example of an environment for executing at least some of the computer code involved in performing the various methods described herein, such as the callback data update module 107. In addition to the callback data update module 107, computing environment 100 also includes, for example, a computer 101, a wide area network (WAN) 102, an end-user equipment (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, computer 101 includes a processor set 110 (including processing circuitry 120 and a cache 121), a communication architecture 111, volatile memory 112, persistent storage device 113 (including an operating system 122 and the callback data update module 107, as identified above), a peripheral device set 114 (including a user interface (UI) device set 123, a storage device 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. Remote server 104 includes a remote database 130. Public cloud 105 includes gateway 140, cloud orchestration module 141, host physical machine set 142, virtual machine set 143, and container set 144.
[0013] Computer 101 can take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or to be developed in the future capable of running programs, accessing networks, or querying databases such as remote database 130. As is well known in the field of computer technology, and depending on that technology, the execution of computer-implemented methods can be distributed among multiple computers and / or multiple locations. On the other hand, in this presentation of computing environment 100, the detailed discussion focuses on a single computer, specifically computer 101, to keep the presentation as simple as possible. Computer 101 can reside in the cloud, even... Figure 1It is not shown in the cloud. On the other hand, computer 101 does not need to be in the cloud, except to the extent that can be definitively indicated.
[0014] Processor set 110 includes one or more computer processors of any type now known or to be developed in the future. Such computer processors, as well as graphics processors, accelerators, coprocessors, etc., are sometimes referred to herein as processing devices. Processing devices and memory operatively coupled to processing devices are sometimes referred to herein as apparatuses. Processing circuitry 120 may be distributed across multiple packages, such as multiple cooperating integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within the processor chip package(s) and is typically used for data or code that should be readily accessible by threads or cores running on processor set 110. Cache memory is typically organized into multiple levels based on its relative proximity to the processing circuitry. Alternatively, some or all of the cache in the processor set may be located “off-chip.” In some computing environments, processor set 110 may be designed to work with qubits and perform quantum computing.
[0015] Computer-readable program instructions are typically loaded onto computer 101 to cause the processor set 110 of computer 101 to perform a series of operational steps to implement a computer-implemented method, such that the instructions thus executed instantiate the method specified in the flowcharts and / or descriptive descriptions of the computer-implemented method included in this document. These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by processor set 110 to control and direct the execution of the computer-implemented method. In computing environment 100, at least some of the instructions for performing the computer-implemented method may be stored in the callback data update module 107 in persistent storage device 113.
[0016] Communication architecture 111 is a signal transmission path that allows various components of computer 101 to communicate with each other. Typically, this architecture consists of switches and conductive paths, such as switches and conductive paths that form buses, bridges, physical input / output ports, etc. Other types of signal communication paths can be used, such as fiber optic communication paths and / or wireless communication paths.
[0017] Volatile memory 112 is any type of volatile memory now known or to be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, but this is not necessary unless explicitly stated otherwise. In computer 101, volatile memory 112 is located in a single package and is internal to computer 101; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located externally relative to computer 101.
[0018] The persistent storage device 113 is any form of non-volatile memory for a computer, now known or to be developed in the future. The non-volatility of this storage device means that the stored data is retained regardless of whether the computer 101 is powered or not, or whether the persistent storage device 113 is directly powered. The persistent storage device 113 may be a read-only memory (ROM), but typically at least a portion of the persistent storage device allows data to be written, deleted, and rewritten. Some common forms of persistent storage devices include hard disks and solid-state storage devices. The operating system 122 may take several forms, such as various known proprietary operating systems employing a kernel, or an operating system with an open-source portable operating system interface type. The code included in the callback data update module 107 typically includes at least some of the computer code involved in performing the computer implementation of the methods described herein.
[0019] Peripheral device set 114 includes the peripheral device set of computer 101. Data communication connections between peripheral devices and other components of computer 101 can be implemented in various ways, such as Bluetooth connections, near field communication (NFC) connections, connections made by cables (such as Universal Serial Bus (USB) type cables), plug-in connections (e.g., secure digital (SD) cards), connections made through local area communication networks, and even connections made through wide area networks such as the Internet. In various embodiments, UI device set 123 may include components such as displays, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage device 124 is an external storage device, such as an external hard drive, or a pluggable storage device, such as an SD card. Storage device 124 may be persistent and / or volatile. In some embodiments, storage device 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 101 requires substantial storage (e.g., where computer 101 locally stores and manages a large database), this storage can be provided by peripheral storage devices designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 125 consists of sensors that can be used in IoT applications. For example, one sensor could be a thermometer, while another could be a motion detector.
[0020] Network module 115 is a collection of computer software, hardware, and firmware that allows computer 101 to communicate with other computers via WAN 102. Network module 115 may include hardware such as a modem or Wi-Fi transceiver, software for packetizing and / or depacketizing data for transmission over the communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control and network forwarding functions of network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software-defined networking (SDN), the control and forwarding functions of network module 115 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for performing computer-implemented methods can typically be downloaded to computer 101 from an external computer or external storage device via a network adapter card or network interface included in network module 115.
[0021] WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances using any technology known now or to be developed in the future for transmitting computer data. In some embodiments, WAN 102 may be replaced by and / or supplemented by a local area network (LAN), which is designed to transmit data between devices located in a local area, such as a Wi-Fi network. WAN and / or LAN typically include computer hardware such as copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.
[0022] End User Equipment (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating computer 101) and can take any of the forms discussed above in conjunction with computer 101. EUD 103 typically receives helpful and useful data from the operation of computer 101. For example, assuming computer 101 is designed to provide recommendations to the end user, these recommendations would typically be transmitted to EUD 103 from network module 115 of computer 101 via WAN 102. In this way, EUD 103 may display or otherwise present the recommendations to the end user. In some embodiments, EUD 103 may be client equipment such as a thin client, heavy client, mainframe, desktop computer, etc.
[0023] Remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 104 can be controlled and used by the same entity operating computer 101. Remote server 104 represents a machine (one or more) that collects and stores helpful and useful data for use by other computers such as computer 101. For example, if computer 101 is designed and programmed to provide recommendations based on historical data, that historical data can be provided to computer 101 from a remote database 130 of remote server 104.
[0024] Public cloud 105 is any computer system that can be used by multiple entities, providing on-demand availability of computer system resources and / or other computing capabilities (especially data storage (cloud storage) and computing power) without the need for direct, active management by users. Cloud computing typically leverages resource sharing to achieve scalability consistency and economy. Direct and active management of the computing resources of public cloud 105 is performed by the computer hardware and / or software of cloud orchestration module 141. The computing resources provided by public cloud 105 are typically implemented by virtual computing environments running on various computers constituting host physical machine set 142, which is the entire domain of physical computers in and / or available to the public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from virtual machine set 143 and / or containers from container set 144. It should be understood that these VCEs can be stored as images and can be transferred between various physical machine hosts as images or after the VCEs are instantiated. Cloud orchestration module 141 manages the transfer and storage of images, deploys new instantiations of VCEs, and manages the active instantiation of VCE deployments. Gateway 140 is a collection of computer software, hardware, and firmware that allow public cloud 105 to communicate via WAN 102.
[0025] Now, we will provide some further explanation of Virtualized Computing Environments (VCEs). A VCE can be stored as an "image." New active instances of a VCE can be instantiated from this image. Two common types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows multiple isolated user-space instances, called containers, to exist. From the perspective of the programs running within them, these isolated user-space instances typically appear as actual computers. Computer programs running on a regular operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running within a container can only use the contents of the container and the devices allocated to the container; this is a characteristic known as containerization.
[0026] Private cloud 106 is similar to public cloud 105, except that computing resources are available only to a single enterprise. While private cloud 106 is depicted as communicating with WAN 102, in other embodiments, private cloud may be completely disconnected from the Internet and accessible only via a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types) typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is bound together by standardization or proprietary technology that enables orchestration, management, and / or data / application portability across the multiple component clouds. In this embodiment, public cloud 105 and private cloud 106 are both part of a larger hybrid cloud.
[0027] Now for reference Figure 2 , Figure 2 Another example computing environment according to various aspects of this disclosure is illustrated. Computing environment 200 includes computing system 202. In a particular embodiment, computing system 202 includes... Figure 1 The computer 101 described herein. The computing system 202 includes one or more services 204 and a callback data update module 206. The one or more services 204 may include one or more services provided by the computing system 202, such as database services. The callback data update module 206 includes a reference code vectorization component 208, a semantic analysis component 210, and a trained model 212. The reference code vectorization component 208 is configured to determine vector representations of one or more reference codes within problem analysis data 216, as described herein with respect to various embodiments. The semantic analysis component 210 is configured to perform semantic analysis, such as natural language processing (NLP), on the problem analysis data 216 to identify data within the problem analysis data 216 that may be related to a specific defect, as described herein with respect to various embodiments. For example, the semantic analysis component 210 may use semantic parsing to identify specific keywords or phrases within the problem analysis data 216 that are related to a specific defect.
[0028] The trained model 212 is configured to be trained with training data using tracked data associated with one or more defects, receiving a vectorized reference code as input and outputting one or more confidence scores associated with the reference code, as described herein with respect to various embodiments. The trained model 212 is also configured to dynamically update the association between the reference code and the data based on the confidence scores. In the event of subsequent occurrences of the same or related defects, the model can provide more relevant problem analysis data to allow for improved defect analysis. As the trained model 212 is continuously trained, its accuracy in identifying relevant data in problem analysis data used to resolve a specific defect is improved. In one embodiment, the callback data update module 206 includes... Figure 1 The callback data update module 107.
[0029] The computing system 202 also includes one or more defect resolvers 214 configured to detect specific defects within the computing system 202 and generate usage data indicating the frequency of access, use, or creation of specific files or other data during the occurrence of the defect and / or attempts to resolve it. In a particular embodiment, the one or more defect resolvers 214 may include software components configured to monitor and detect specific defects. In other embodiments, the defect resolver may include a user performing a defect analysis and resolution function. The computing system 202 is also configured to generate a debug data file 218 based on a portion of fault analysis data 216 identified by the trained model 212 as useful for resolving defects.
[0030] The computing environment 200 also includes a server 222 that communicates with the computing system 202 via a network 220. In a particular embodiment, the server 222 is associated with a service provider of one or more services 204. In one or more embodiments, the computing system 202 is configured to send debug data files 218 to the server 222 for analysis to determine the cause of defects.
[0031] Now for reference Figure 3 This illustrates the use of various aspects of this disclosure for Figure 2 An example of a processing flow in a computing environment. In processing flow 300, computing system 202 identifies error (error 2) from a plurality of possible errors (e.g., error 1, error 2, error 3) 302 as a causing error of a defect in computing system 202. In a particular embodiment, the error is a software or hardware error associated with a service provided by computing system 202. Computing system 202 collects 306 data related to error 2 from a corpus of problem analysis data 216. In a particular embodiment, computing system 202 collects the portion of the data related to error 2 by semantic parsing 308 and collects another portion of the data using problem analysis vector association 310. For example, semantic parsing 308 may be used to identify predetermined words or phrases in problem analysis data 216 that are particularly related to a defect or error. In another example, problem analysis vector association 310 may be used to vectorize reference code as described herein with respect to various embodiments.
[0032] The computing system 202 may optionally take pre-emptive action 312 by initiating a callback to server 222. The computing system 202 tracks 314 the use of defects by the defect resolver to determine the frequency of access to or creation of defect-related data, such as the access to or creation of specific files or other data associated with defects, in relation to the identification and / or resolution of various defects. The computing system 312 identifies 316 useful data for training the model by determining, based on data usage associated with one or more defects, to be used as part of the problem analysis data 216 as training data for the model.
[0033] The computing system 202 creates a 318-trained model, which is configured to be trained with training data using tracked data associated with one or more defects. The trained model is configured to receive 320 vectorized reference codes as input and output one or more data confidence scores associated with the reference codes, as described herein with respect to various embodiments.
[0034] Now for reference Figure 4 , Figure 4 A flowchart illustrating an example process for model-based updates of data rollback according to various aspects of this disclosure is shown. Computing system 202 receives 402 problem analysis data associated with the computing system. In one or more embodiments, an initiation is made by determining that a defect has occurred in computing system 202 (such as a hardware or software defect in the services provided by computing system 202). Figure 4 The processing. In a particular embodiment, the problem analysis data includes data files.
[0035] The computing system 202 identifies one or more reference codes within the problem analysis data that are associated with one or more defects. In a particular embodiment, each of the one or more reference codes indicates a diagnostic result of the computing system 202 associated with the one or more defects. Based on the data used in relation to the one or more defects, the computing system 202 determines, 406, a portion of the problem analysis data to be used as training data for the model. In a particular embodiment, the computing system 202 further determines the portion of the problem analysis data to be used as training data for the model by performing semantic parsing on the problem analysis data to determine the portion of the problem analysis data associated with the defects.
[0036] The computing system 202 inputs a representation of each of one or more reference codes within a portion of the problem analysis data into the model. The model is configured to output one or more data confidence scores based on the one or more reference codes. In a particular embodiment, the representation of each of the one or more reference codes includes a vector-based representation. The computing system 202 updates the association between the one or more reference codes and one or more defects based on the one or more data confidence scores.
[0037] Now for reference Figure 5 , Figure 5 A flowchart is shown for another example of a process for updating callback data limits and priorities in accordance with various aspects of this disclosure. Figure 5 Example processing includes about Figure 4 The example processing steps described herein, and also include determining, based on data usage associated with one or more defects, 406 a portion of the problem analysis data to be used as training data for the model, including tracking 502 data usage associated with one or more defects to determine a portion of the problem analysis data to be used as training data for the model. In a particular embodiment, tracking data usage associated with one or more defects includes using a heatmap to track data usage associated with one or more defects.
[0038] exist Figure 5 In the example processing, inputting the representation of each of one or more reference codes within a portion of the problem analysis data into the model includes generating a vector-based representation of each of the one or more reference codes. In a particular embodiment, the vector-based representation of each of the one or more reference codes is generated based on one or more features found in a corpus of the problem analysis data. In a particular embodiment, the features used for vectorization of the reference codes may include one or more fields in the problem analysis data, related defect links (e.g., links created between related defects in the problem analysis data), and / or sentence embeddings related to textual information found in the problem analysis data.
[0039] Figure 5The example processing also includes computing system 202 determining, 506, a subset of problem analysis data to be included in a debug data file associated with the corresponding reference code, based on one or more data confidence scores. Computing system 202 stores the subset of problem analysis data, 508, in the debug data file and sends the debug data file, 510, to a server. In a particular embodiment, the server determines a solution to the defect based on the debug data file and communicates the solution to computing system 202. In another particular embodiment, one or more users associated with a service provider utilize the debug data file to determine a solution to the defect. One or more users associated with the service provider can communicate the solution to computing system 202 and / or users associated with computing system 202, or perform actions to resolve the defect.
[0040] Various aspects of this disclosure are described through narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). With respect to any flowchart, depending on the technology involved, operations may be performed in a different order than those shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.
[0041] Computer Program Product Embodiment (“CPP Embodiment” or “CPP”) is a term used in this disclosure to describe any group of one or more storage media (also referred to as “media”) commonly included in a group of one or more storage devices, said group of one or more storage devices commonly including machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A “storage device” is any tangible device capable of holding and storing instructions for use by a computer processor. Without limitation, a computer-readable storage medium can be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include: magnetic disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), optical disc read-only memory (CD-ROM), digital universal disc (DVD), memory sticks, floppy disks, mechanical encoding devices (such as punched cards or pits / ridges formed on the main surface of the disk), or any suitable combination of the foregoing. As used in this disclosure, computer-readable storage media should not be construed as storing transient signals, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses through fiber optic cables, electrical signals transmitted through wires, and / or other transmission media. As those skilled in the art will understand, data typically moves at some occasional points in time during normal operation of the storage device, such as during access, defragmentation, or garbage collection, but this does not make the storage device transient, because the data is not transient when it is stored.
[0042] The description of various embodiments in this disclosure is for illustrative purposes only and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles of the embodiments, their practical application, or technical improvements to existing technologies in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method comprising: Receive problem analysis data associated with the computing system; Identify one or more reference codes within the problem analysis data that are associated with one or more defects; Based on the data associated with the one or more defects, a portion of the problem analysis data is determined to be used as training data for the model; The model is input with a representation of each of one or more reference codes within the portion of the problem analysis data, and the model is configured to output one or more data confidence scores based on the one or more reference codes. as well as The association between the one or more reference codes and the one or more defects is updated based on the one or more data confidence scores.
2. The method of claim 1, wherein determining a portion of the problem analysis data to be used as training data for the model based on data associated with the one or more defects comprises: The data associated with the one or more defects is used to determine the portion of the problem analysis data to be used as training data for the model.
3. The method of claim 2, wherein tracking data associated with the one or more defects uses: Heatmaps are used to track data associated with the one or more defects.
4. The method of claim 1, wherein determining the problem analysis data to be used as training data for the model further comprises: Semantic parsing is performed on the problem analysis data to identify a subset of the problem analysis data that is relevant to the defects.
5. The method of claim 1, wherein the representation of each of the one or more reference codes comprises a vector-based representation.
6. The method of claim 1, wherein each of the one or more reference codes indicates a diagnostic result of the computing system associated with the one or more defects.
7. The method of claim 1, wherein a subset of the problem analysis data is determined based on the one or more data confidence scores to be included in a debug data file associated with the corresponding reference code.
8. The method according to claim 7, further comprising: Store the problem analysis data in the debug data file.
9. The method of claim 1, wherein the defect includes a software defect in the computing system.
10. The method according to claim 1, wherein the problem analysis data includes a data file.
11. The method of claim 7, further comprising: Send the debug data file to the server.
12. An apparatus comprising: Processing equipment; as well as A memory operatively coupled to a processing device, wherein the memory stores computer program instructions that, when executed, cause the processing device to: Receive problem analysis data associated with the computing system; Identify one or more reference codes within the problem analysis data that are associated with one or more defects; Based on the data associated with the one or more defects, a portion of the problem analysis data is determined to be used as training data for the model; The model is input with a representation of each of one or more reference codes within the portion of the problem analysis data, and the model is configured to output one or more data confidence scores based on the one or more reference codes. as well as The association between the one or more reference codes and the one or more defects is updated based on the one or more data confidence scores.
13. The device of claim 12, wherein using data associated with the one or more defects to determine a portion of the problem analysis data to be used as training data for the model comprises: The data associated with the one or more defects is used to determine the portion of the problem analysis data to be used as training data for the model.
14. The apparatus of claim 12, wherein the problem analysis data determined to be used as training data for the model further comprises: Semantic parsing is performed on the problem analysis data to identify a subset of the problem analysis data that is relevant to the defects.
15. The device of claim 12, wherein the representation of each of the one or more reference codes comprises a vector-based representation.
16. The device of claim 12, wherein each of the one or more reference codes indicates a diagnostic result of the computing system associated with the one or more defects.
17. A computer program product comprising a computer-readable storage medium, wherein the computer-readable storage medium includes computer program instructions that, when executed: Receive problem analysis data associated with the computing system; Identify one or more reference codes within the problem analysis data that are associated with one or more defects; Based on the data associated with the one or more defects, a portion of the problem analysis data is determined to be used as training data for the model; The model is input with a representation of each of one or more reference codes within the portion of the problem analysis data, and the model is configured to output one or more data confidence scores based on the one or more reference codes. as well as The association between the one or more reference codes and the one or more defects is updated based on the one or more data confidence scores.
18. The computer program product of claim 17, wherein determining a portion of the problem analysis data to be used as training data for the model based on data associated with the one or more defects comprises: The data associated with the one or more defects is used to determine the portion of the problem analysis data to be used as training data for the model.
19. The computer program product of claim 17, wherein the problem analysis data determined to be used as training data for the model further comprises: Semantic parsing is performed on the problem analysis data to identify a subset of the problem analysis data that is relevant to the defects.
20. The computer program product of claim 17, wherein the representation of each of the one or more reference codes comprises a vector-based representation.