A data processing method, apparatus, device, medium, and product

CN119066098BActive Publication Date: 2026-09-25SHANGHAI GUBO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411181710.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-27
Publication Date
2026-09-25
Estimated Expiration
2044-08-27

AI Technical Summary

Technical Problem

半导体测试数据的数据量较大,因此,会占用过多的存储空间

Benefits of technology

[0019]根据本发明的另一方面,提供了一种计算机程序产品,所述计算机程序在被处理器执行时实现如本发明实施例中任一所述的数据处理方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119066098B_ABST
    Figure CN119066098B_ABST
Patent Text Reader

Abstract

A data processing method, device, equipment, medium and product are disclosed. The method comprises: obtaining a data set, at least one original test item, at least one calculation test item composed of original test items, and at least one split test item composed of original test items, wherein the data set comprises at least one wafer identifier; determining a plurality of data processing tasks and a dependency relationship between the data processing tasks according to the at least one wafer identifier, the at least one original test item, the at least one calculation test item composed of original test items, and the at least one split test item composed of original test items; and executing the data processing tasks in sequence according to the dependency relationship between the data processing tasks and the data set to obtain data processing results corresponding to the data processing tasks. The technical solution disclosed in the application can reduce the amount of occupied storage space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of semiconductor technology, and in particular to a data processing method, apparatus, device, medium and product. Background Technology

[0002] Semiconductor test data is typically large in scale and yields a wide variety of results. Analysis involves not only processing the raw test data but also calculating statistical indicators, industry-standard parameters, and more. Existing data processing solutions include real-time analysis and offline analysis.

[0003] Real-time analysis performs calculations as needed, avoiding large-scale data storage. However, real-time analysis is susceptible to the volume of data; with large datasets, processing efficiency is low, making real-time analysis impossible. For semiconductor testing data, when analyzing a single chip on a wafer, the data volume typically reaches tens of billions, making real-time analysis unsuitable.

[0004] Offline analysis involves performing calculations and caching on the raw test data offline. Semiconductor test data is large in volume, thus consuming excessive storage space. Summary of the Invention

[0005] This invention provides a data processing method, apparatus, device, medium, and product that can reduce the amount of storage space required.

[0006] According to one aspect of the present invention, a data processing method is provided, comprising:

[0007] The dataset comprises at least one original test item, at least one computational test item consisting of the original test items, and at least one split test item consisting of the original test items, wherein the dataset includes at least one wafer identifier;

[0008] Based on at least one wafer identifier, at least one original test item, at least one computational test item consisting of the original test items, and at least one split test item consisting of the original test items, determine multiple data processing tasks and the dependencies between the data processing tasks.

[0009] Based on the dependencies between the data processing tasks and the dataset, each data processing task is executed sequentially to obtain the data processing results corresponding to each data processing task.

[0010] According to another aspect of the present invention, a data processing apparatus is provided, the data processing apparatus comprising:

[0011] An acquisition module is used to acquire a dataset, at least one original test item, at least one computational test item consisting of the original test items, and at least one split test item consisting of the original test items, wherein the dataset includes at least one wafer identifier;

[0012] The task generation module is used to determine multiple data processing tasks and the dependencies between them based on at least one wafer identifier, at least one original test item, at least one computational test item composed of the original test items, and at least one split test item composed of the original test items.

[0013] The task execution module is used to execute each data processing task sequentially according to the dependencies between the data processing tasks and the dataset, and to obtain the data processing results corresponding to each data processing task.

[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data processing method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data processing method described in any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer program product is provided, which, when executed by a processor, implements the data processing method as described in any of the embodiments of the present invention.

[0020] This invention provides an embodiment that obtains a dataset, at least one original test item, at least one computational test item composed of the original test items, and at least one split test item composed of the original test items, wherein the dataset includes at least one wafer identifier; based on the at least one wafer identifier, at least one original test item, at least one computational test item composed of the original test items, and at least one split test item composed of the original test items, multiple data processing tasks and the dependencies between each data processing task are determined; based on the dependencies between the data processing tasks and the dataset, each data processing task is executed sequentially to obtain the data processing result corresponding to each data processing task, thereby reducing the amount of storage space occupied.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a data processing method according to an embodiment of the present invention;

[0024] Figure 2 This is a flowchart of another data processing method in an embodiment of the present invention;

[0025] Figure 3 This is a schematic diagram of the structure of a data processing device according to an embodiment of the present invention;

[0026] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0030] Example 1

[0031] Figure 1 This is a flowchart illustrating a data processing method provided in an embodiment of the present invention. This embodiment is applicable to data processing situations. The method can be executed by the data processing device in this embodiment of the present invention, which can be implemented in software and / or hardware, such as... Figure 1 As shown, the method specifically includes the following steps:

[0032] S110, obtain a dataset, at least one original test item, at least one computational test item consisting of the original test items, and at least one split test item consisting of the original test items.

[0033] The dataset includes at least one wafer identifier. The dataset also includes semiconductor test results corresponding to at least one wafer identifier. For example, the dataset may include semiconductor test results for multiple test items corresponding to wafer identifier A, semiconductor test results for multiple test items corresponding to wafer identifier B, and semiconductor test results for multiple test items corresponding to wafer identifier C.

[0034] Wherein, the at least one original test item can be: at least one original test item identifier, which can be the name of the original test item. The calculation test item composed of the original test items is used to characterize: performing calculations on multiple original test items. For example, the calculation test item can be: the name of original test item 1 + the name of original test item 2. The split test item composed of the original test items is used to characterize: splitting multiple original test items. For example, the split test item can be: split(Test Item,':',2)as num.

[0035] The dataset is used to represent a predefined data range. Specifically, the dataset can be obtained by querying the total semiconductor test data (test results) based on the key information input by the user, thus obtaining the dataset corresponding to the key information.

[0036] Optionally, obtain the dataset, including:

[0037] The main interface is displayed in response to a preset trigger action.

[0038] In response to a trigger action on the New Dataset control in the main interface, display the dataset selection interface;

[0039] In response to a trigger operation on the preview control in the dataset selection interface, the dataset corresponding to the test attribute information input by the user is obtained and the dataset is displayed in the dataset selection interface. The test attribute information includes at least one of the following: test stage, test time, batch number, test batch, and wafer identifier.

[0040] The main interface may include: a dataset list, a batch delete control, a filter list control, and a new dataset control. The dataset list includes: dataset name, data table source, testing phase, creator, creation time, modifier, modification time, and corresponding operation controls (operation controls may include: enter icon analysis control, dataset settings control, delete control, and more controls).

[0041] The dataset selection interface includes a data range selection area and a data selection overview area. The data range selection area includes a preview control, a reset control, and a test attribute information input area. The data selection overview area displays detailed dataset information (including: test stage, batch number, test batch, wafer group identifier, wafer identifier, total number of chips, number of chips that passed testing, test pass rate, etc.). Users can directly enter the test attribute information in the test attribute information input area, or select the test attribute information from a drop-down list.

[0042] It should be noted that a batch part number contains multiple wafer groups, a wafer group contains multiple wafers, and a wafer contains multiple chips. The smallest granularity for defining a dataset is the wafer.

[0043] Optionally, obtain at least one original test item, including:

[0044] In response to a trigger operation on the next control in the dataset selection interface, a test item selection interface is displayed, wherein the test item selection interface includes a query area and a test item list display area;

[0045] In response to a trigger operation on a control in the query area, the list of test items in the test item list display area is updated according to the query information entered by the user;

[0046] In response to a selection operation for a test item in the test item list, the selected test item is identified as the original test item.

[0047] The query area includes a query information editing area. The test item list includes a test item name and a test item type.

[0048] It should be noted that the test item query is based on the test item name. The query area may also include: the test item selection method, for example, it can be customized (directly selecting the required test items from the overall list), or it can be obtained by entering a regular expression (filtering the list with a regular expression to obtain the required test items. For example, it can be: test items whose names do not contain XXXX, or test items whose names contain XXX). This embodiment of the invention does not impose any restrictions on this.

[0049] It should be noted that after determining the dataset, i.e., defining the wafer range, the test item range is further defined. Each chip on each wafer may undergo multiple tests. During data processing, only a subset of test items can be selected. Optionally, at least one computational test item consisting of the original test items is obtained, including:

[0050] In response to the triggering operation of the next step control in the test item selection interface, the test item setting interface is displayed;

[0051] In response to a triggering operation of adding a new calculation test item control in the calculation test item settings interface, a calculation test item editing interface is displayed, wherein the calculation test item editing interface includes: a first editing area, a first test item list display area, and a first function display area;

[0052] In response to a trigger operation on a test item in the first test item list display area, the triggered test item is displayed in the first editing area; in response to a trigger operation on a function in the first function display area, the triggered function is displayed in the first editing area.

[0053] In response to a trigger operation on the confirmation control in the calculation test item editing interface, a calculation test item is generated based on the triggered test item and the triggered function displayed in the first editing area.

[0054] The calculation test item setting interface includes: a calculation test item list display area and a control for adding a new calculation test item. The calculation test item editing interface includes: a calculation test item name editing area, a first editing area, a first test item list display area, and a first function display area.

[0055] The first test item list display area is used to display the list of test items selected in the previous step. The first function display area is used to display the list of calculation formulas, for example, as shown in the first column of Table 1 (the second column of Table 1 is used to explain the meaning of each calculation formula):

[0056] Table 1

[0057]

[0058]

[0059] Specifically, in response to a trigger operation on a test item in the first test item list display area, the triggered test item is displayed in the first editing area; in response to a trigger operation on a function in the first function display area, the triggered function is displayed in the first editing area. This can be achieved by the user double-clicking a test item in the first test item list, triggering the test item, and then displaying the triggered test item in the first editing area. After the function is triggered, it is displayed after the previously triggered test item. For example, if the user double-clicks test item 1 in the first test item list, then double-clicks the summation function, and then double-clicks test item 2 in the first test item list, then the first editing area will display: Test item 1 + Test item 2.

[0060] Specifically, the method for generating calculation test items based on the triggered test items and functions displayed in the first editing area can be as follows: The triggered test items and functions are sequentially concatenated according to the user's triggering order to obtain the calculation test items. For example, if the user double-clicks test item 1 in the first test item list, then double-clicks the summation function, and then double-clicks test item 2 in the first test item list, the calculation test item would be: Test item 1 + Test item 2.

[0061] Optionally, obtain at least one split test item consisting of the original test item, including:

[0062] In response to the triggering operation of the next step control in the calculation test item settings interface, the split test item settings interface is displayed;

[0063] In response to the triggering operation of adding a new split test item control in the split test item settings interface, a split test item editing interface is displayed, wherein the split test item editing interface includes: a second editing area, a second test item list display area, and a second function display area;

[0064] In response to a trigger operation on a test item in the second test item list display area, the triggered test item is displayed in the second editing area; in response to a trigger operation on a function in the second function display area, the triggered function is displayed in the editing area.

[0065] In response to a trigger operation on the confirm control in the split test item editing interface, a split test item is generated based on the triggered test item and the triggered function displayed in the second editing area.

[0066] The test item splitting settings interface includes: a test item splitting list display area and a control for adding test items splitting the test item. The test item splitting editing interface includes: a test item splitting name editing area, a second editing area, a second test item list display area, and a second function display area.

[0067] The second test item list display area is used to display the list of test items selected in the previous step. The second function display area is used to display the list of calculation formulas, for example, as shown in the first column of Table 2 (the second column of Table 2 is used to explain the meaning of each calculation formula):

[0068] Table 2

[0069]

[0070]

[0071] It should be noted that if the test item name is 123:test, after calculation using formula 1, the test item name will be changed to test. After calculation using formula 2, 123 will be used as an attribute of this test item. Some companies may name test items as: Voltage Test: 5V, Voltage Test: 10V, etc. After splitting the test item, the name can be unified as: Voltage Test, with attributes of 5V and 10V respectively.

[0072] S120, based on at least one wafer identifier, at least one original test item, at least one computational test item consisting of the original test items, and at least one split test item consisting of the original test items, determine multiple data processing tasks and the dependencies between the data processing tasks.

[0073] The data processing tasks may include: a first type of data processing task and a second type of data processing task. The dependency relationship between the data processing tasks may be that the first type of data processing task is a predecessor task of the second type of data processing task.

[0074] Specifically, the method for determining multiple data processing tasks and the dependencies between them based on at least one wafer identifier, at least one original test item, at least one computational test item composed of the original test items, and at least one split test item composed of the original test items can be as follows: Multiple first-type data processing tasks are generated based on each computational test item and at least one wafer identifier, wherein the total number of first-type data processing tasks is the same as the number of computational test items; multiple second-type data processing tasks are generated based on each split test item and at least one wafer identifier, wherein the total number of second-type data processing tasks is the same as the number of split test items, and the first-type data processing tasks are precursor tasks of the second-type data processing tasks.

[0075] Optionally, based on at least one wafer identifier, at least one original test item, at least one computational test item consisting of the original test items, and at least one split test item consisting of the original test items, multiple data processing tasks and the dependencies between the data processing tasks are determined, including:

[0076] Based on each computational test item and at least one wafer identifier, multiple first-type data processing tasks are generated, wherein the total number of first-type data processing tasks is the same as the number of computational test items;

[0077] Based on each split test item and at least one wafer identifier, multiple second-type data processing tasks are generated, wherein the total number of second-type data processing tasks is the same as the number of split test items, and the first-type data processing tasks are the precursor tasks of the second-type data processing tasks.

[0078] It should be noted that the first type of data processing task is a precursor task to the second type of data processing task. That is, the first type of data processing task needs to be processed in advance, and the second type of data processing task needs to use the processing result of the first type of data processing task when it is processing.

[0079] Specifically, based on each computational test item and at least one wafer identifier, multiple first-type data processing tasks are generated. This can be done in the following way: one first-type data processing task is generated for each computational test item, and each first-type data processing task is a data processing task for all selected wafer identifiers.

[0080] Specifically, based on each split test item and at least one wafer identifier, multiple second-type data processing tasks are generated. This can be done in the following way: one second-type data processing task is generated for each split test item, and each second-type data processing task is a data processing task for all selected wafer identifiers.

[0081] Optionally, based on the dependencies between the data processing tasks and the dataset, each data processing task is executed sequentially to obtain the data processing results corresponding to each data processing task, including:

[0082] After performing the first type of data processing task, perform the second type of data processing task to obtain a list of processing results;

[0083] The execution of the first type of data processing task includes: obtaining each original test item contained in the calculation test item corresponding to the first type of data processing task; querying the dataset to obtain the test results of each original test item corresponding to the wafer identifier; performing a first calculation on the test results of each original test item corresponding to the wafer identifier according to the function in the calculation test item to obtain the processing result; and generating a first list based on the processing result.

[0084] Executing the second type of data processing task includes: obtaining each original test item contained in the split test item corresponding to the second type of data processing task; performing a second calculation on the identification information of each original test item according to the function in the split test item to obtain attribute information; adding the attribute information to the first list to obtain a processing result list.

[0085] The identification information of the original test item can be the name of the original test item, and the identification information of the original test item carries attribute information. For example, the name of the original test item can be: Voltage Test: 5V. The 5V in the name of the original test item is attribute information.

[0086] It should be noted that, during the execution of the first type of data processing task, the method for querying the dataset to obtain the test results of each original test item corresponding to the wafer identifier can be as follows: based on the wafer identifier and the test item name, query the dataset to obtain the test results of each original test item corresponding to the wafer identifier.

[0087] For example, the computational test items corresponding to the first type of data processing task may include: original test item 1 and original test item 2. Selected wafers are identified as wafer A and wafer B. The dataset is queried to obtain the test results of wafer A when performing original test item 1, wafer B when performing original test item 1, wafer A when performing original test item 2, and wafer B when performing original test item 2. Based on the test results of wafer A and wafer B when performing original test item 1, the test results of wafer A and wafer B when performing original test item 2, and the function in the computational test items, a first calculation is performed to obtain the processing result. A first list is then generated based on the processing result.

[0088] It should be noted that during the execution of the second type of data processing task, the identification information of the test items is split, and the resulting information is determined as attributes. For example, if the identification information of a test item is its name, and the test item names are: Voltage Test: 5V, Voltage Test: 10V, then after splitting the test item name, the test item name can be unified as: Voltage Test, with attribute information of 5V and 10V respectively.

[0089] S130, based on the dependencies between the data processing tasks and the dataset, execute each data processing task in sequence to obtain the data processing results corresponding to each data processing task.

[0090] Specifically, the method for sequentially executing each data processing task based on the dependencies between the data processing tasks and the dataset to obtain the data processing results corresponding to each data processing task can be as follows: Slice each data processing task to obtain multiple data processing task slices; determine the dependencies between the multiple data processing task slices based on the dependencies between the data processing tasks; sequentially execute each data processing task slice based on the dependencies between the data processing task slices and the dataset to obtain the data processing results corresponding to each data processing task slice; and determine the data processing result corresponding to each data processing task based on the data processing results corresponding to each data processing task slice and the dependencies between the data processing task slices. Alternatively, the method for sequentially executing each data processing task based on the dependencies between the data processing tasks and the dataset to obtain the data processing results corresponding to each data processing task can be as follows: Submit each data processing task to a message queue, wait for the executor to retrieve the message, and proceed with the next step of processing. Multiple executors can be deployed on different machines. The executor listens to the message queue, and when it detects a task to be processed, it retrieves the message and processes it.

[0091] Optionally, based on the dependencies between the data processing tasks and the dataset, each data processing task is executed sequentially to obtain the data processing results corresponding to each data processing task, including:

[0092] Each data processing task is sliced ​​based on the total number of wafer identifiers to obtain multiple data processing task slices and the dependencies between each data processing task slice, wherein the total number of data processing task slices is the same as the total number of wafer identifiers.

[0093] Based on the dependencies of each data processing task slice and the dataset, each data processing task slice is executed sequentially to obtain the data processing results corresponding to each data processing task slice.

[0094] Based on the data processing results corresponding to each data processing task slice and the dependencies between each data processing task slice, the data processing results corresponding to each data processing task are determined.

[0095] The dependencies between the data processing task slices can be determined by the data processing task identifier carried by each data processing task slice. For example, if the data processing task slices carry the same data processing task identifier, that is, all data processing task slices belong to the same data processing task. If the data processing task slices carry different data processing task identifiers, then the dependencies between data processing task slices carrying the same data processing task identifier are determined to belong to the same data processing task.

[0096] Specifically, the method for slicing each data processing task based on the total number of wafer identifiers to obtain multiple data processing task slices and the dependencies between them can be as follows: If the total number of wafer identifiers is M, then slicing each data processing task will yield M data processing task slices. The dependencies between data processing task slices belonging to the same data processing task are defined as belonging to the same data processing task. The dependencies between data processing task slices belonging to different data processing tasks are the same as the dependencies between data processing tasks, where M is a positive integer.

[0097] In a specific example, if there are a total of 4 wafer identifiers, the data processing tasks include: data processing task R and data processing task T. Data processing task R is the predecessor task of data processing task T. Data processing task R is sliced ​​to obtain 4 data processing task slices r. Data processing task T is sliced ​​to obtain 4 data processing task slices t. The dependency relationship between each data processing task slice r is that they belong to the same data processing task. The dependency relationship between each data processing task slice t is the same as the dependency relationship between data processing task R and data processing task T. That is, data processing task slice r is the predecessor task of data processing task t.

[0098] Specifically, based on the dependencies of each data processing task slice and the dataset, the data processing task slices are executed sequentially to obtain the data processing results corresponding to each data processing task slice. The method is as follows: if each data processing task slice belongs to the same data processing task, the data processing task slices are executed in parallel. If each data processing task slice belongs to different data processing tasks, the data processing task slices belonging to the predecessor task are executed first according to the dataset, and then the data processing task slices belonging to the successor task are executed according to the dataset to obtain the data processing results corresponding to each data processing task slice.

[0099] Specifically, the method for determining the data processing results corresponding to each data processing task slice based on the data processing results corresponding to each data processing task slice and the dependencies between each data processing task slice can be as follows: After obtaining the data processing results corresponding to each data processing task slice, the data processing results corresponding to each data processing task slice are concatenated according to the dependencies between each data processing task slice to obtain the data processing results corresponding to each data processing task.

[0100] In a specific example, if there are a total of 4 wafer identifiers, the data processing tasks include: data processing task R and data processing task T. Data processing task R is a precursor task to data processing task T. Data processing task R is sliced ​​to obtain 4 data processing task slices r. Data processing task T is sliced ​​to obtain 4 data processing task slices t. First, the 4 data processing task slices r are executed in parallel to obtain the data processing results corresponding to the 4 data processing task slices r. Then, the 4 data processing task slices t are executed in parallel to obtain the data processing results corresponding to the 4 data processing task slices t. The data processing results corresponding to the 4 data processing task slices r are concatenated to obtain the data processing result corresponding to data processing task R. Finally, the data processing results corresponding to the 4 data processing task slices t are concatenated to obtain the data processing result corresponding to data processing task T.

[0101] It should be noted that when data processing tasks are sliced, the resulting data processing task slices belonging to the same data processing task can be executed in parallel. However, if two data processing task slices belong to different data processing tasks and the two data processing tasks have a dependency relationship, then the two data processing task slices need to be executed serially.

[0102] Optionally, the data processing tasks include: a first type of data processing task and a second type of data processing task;

[0103] Based on the total number of wafer identifiers, each data processing task is sliced ​​to obtain multiple data processing task slices and the dependencies between each data processing task slice, including:

[0104] Based on the total number of wafer identifiers, each first type of data processing task is sliced ​​to obtain multiple data processing task slices corresponding to the first type of data processing task.

[0105] Based on the total number of wafer identifiers, each second type of data processing task is sliced ​​to obtain multiple data processing task slices corresponding to the second type of data processing task.

[0106] Based on the dependency relationship between the first type of data processing task and the second type of data processing task, determine the dependency relationship between multiple data processing task slices corresponding to the first type of data processing task and multiple data processing task slices corresponding to the second type of data processing task.

[0107] It should be noted that when slicing the first type of data processing task, the number of data processing task slices obtained is the same as the total number of wafer identifiers. Similarly, when slicing the second type of data processing task, the number of data processing task slices obtained is the same as the total number of wafer identifiers. Specifically, based on the dependency relationship between the first and second type of data processing tasks, the dependency relationship between multiple data processing task slices corresponding to the first type of data processing task and multiple data processing task slices corresponding to the second type of data processing task can be determined as follows: if the first data processing task slice belongs to the first type of data processing task and the second data processing task slice belongs to the second type of data processing task, then the dependency relationship between the first and second data processing task slices is determined as the first data processing task slice being the predecessor task of the second data processing task slice.

[0108] Optionally, based on the dependencies between the data processing tasks and the dataset, each data processing task is executed sequentially to obtain the data processing results corresponding to each data processing task, including:

[0109] If all the preceding tasks of the current data processing task have been completed, then the current data processing task is executed based on the dataset to obtain the data processing result corresponding to the current data processing task. The preceding tasks are the data processing tasks that the current data processing task depends on.

[0110] In a specific example, such as Figure 2 As shown in the figure, this embodiment of the invention provides a task scheduling system, which includes a task scheduler, a slicer, an executor, and a task monitor. The task scheduler determines which data processing task to execute first based on the dependencies between data processing tasks (executing the predecessor task first, then the successor task). After the predecessor task is completed, the task scheduler automatically triggers the execution of the successor task. For example, it first checks whether the predecessor task of the task to be processed has been successfully executed; if so, it executes the task to be processed.

[0111] After the task scheduler starts a task, the slicer first processes the data (e.g., ...) data. Figure 2 The slicer divides the data into n tasks. For example, it slices the data based on the total number of wafer identifiers, v, resulting in v task slices. The slicer ensures that the result of summing the individual slices after slicing is equivalent to the result obtained by processing the entire task according to the original data range.

[0112] To improve system concurrency, this embodiment of the invention uses a message queue to decouple the executor. Task slices, generated by the slicer, are uniformly submitted to the message queue, waiting for the executor to retrieve the message and proceed with further processing.

[0113] The executor is the data processing module in this embodiment of the invention. Due to the use of message queues for decoupling, multiple executors can be deployed on different machines, allowing for flexible control of processing concurrency and management of computing power. In this embodiment, V executors can be deployed, each processing one of V task slices. The executor listens to the message queue, retrieving and processing any pending slices. The slicer stores the processed data and unprocessed test results in a cache table. Each data processing task has a unique cache table, which is shared by all executors.

[0114] The task monitor continuously monitors the health of data processing tasks in real time. If a task times out or fails, it will be retried a limited number of times. After exceeding the maximum number of retries, the task will enter a final failure state. In a group of data processing tasks, once one task enters the final failure state, the entire group of data processing tasks is considered to have failed. The task monitor, as a compensation mechanism in this embodiment of the invention, can improve the stability of the system.

[0115] In a specific example, the data processing method provided by this embodiment of the invention includes the following steps:

[0116] 1. Define the data range and design the processing scheme (computational test items consisting of original test items and / or split test items consisting of original test items).

[0117] Users first need to define the range of data to be processed. This embodiment of the invention does not impose strict limitations on the amount of data and requires the design of specific processing schemes. This embodiment of the invention supports operations such as splitting, modifying, and combining the original data to obtain new indicators. There are predecessor and successor relationships between operations, and successor tasks can further process the results of predecessor tasks.

[0118] 2. Submit the task

[0119] The data range and processing scheme will be converted into a set of processing tasks and submitted to the system to which this embodiment of the invention belongs. The task scheduler will determine which task to execute first. After the predecessor task is successfully executed, the successor task will be automatically scheduled to be executed.

[0120] 3. Data range slicing and distribution

[0121] The slicer divides the task into slices and distributes them using a message queue.

[0122] 4. Perform slicing

[0123] Several executors will listen to the message queue and begin specific data processing upon receiving a message.

[0124] This invention improves computational parallelism through a distributed system and incorporates a monitoring program as a compensation mechanism to enhance system robustness. Processing results are stored in a database cache table for easier subsequent use. This invention enables the processing and storage of massive amounts of test results, thus promoting semiconductor data analysis and production decision-making.

[0125] The technical solution of this embodiment obtains a dataset, at least one original test item, at least one computational test item composed of the original test items, and at least one split test item composed of the original test items, wherein the dataset includes at least one wafer identifier; based on at least one wafer identifier, at least one original test item, at least one computational test item composed of the original test items, and at least one split test item composed of the original test items, multiple data processing tasks and the dependencies between each data processing task are determined; based on the dependencies between each data processing task and the dataset, each data processing task is executed sequentially to obtain the data processing result corresponding to each data processing task, thereby reducing the amount of storage space occupied.

[0126] Example 2

[0127] Figure 3 This is a schematic diagram of a data processing device provided in an embodiment of the present invention. This embodiment is applicable to data processing applications. The device can be implemented using software and / or hardware, and can be integrated into any device that provides data processing functionality, such as… Figure 3 As shown, the data processing device specifically includes: an acquisition module 310, a task generation module 320, and a task execution module 330.

[0128] The acquisition module is used to acquire a dataset, at least one original test item, at least one computational test item composed of the original test items, and at least one split test item composed of the original test items, wherein the dataset includes at least one wafer identifier;

[0129] The task generation module is used to determine multiple data processing tasks and the dependencies between them based on at least one wafer identifier, at least one original test item, at least one computational test item composed of the original test items, and at least one split test item composed of the original test items.

[0130] The task execution module is used to execute each data processing task sequentially according to the dependencies between the data processing tasks and the dataset, and to obtain the data processing results corresponding to each data processing task.

[0131] The above-described products can perform the methods provided in any embodiment of the present invention, and have the corresponding functional modules and beneficial effects for performing the methods.

[0132] Example 3

[0133] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0134] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0135] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0136] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data processing methods.

[0137] In some embodiments, the data processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data processing method by any other suitable means (e.g., by means of firmware).

[0138] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0139] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0140] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0141] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0142] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0143] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0144] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0145] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the data processing method according to any embodiment of the invention.

[0146] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0147] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A data processing method, characterized in that, include: The dataset comprises at least one original test item, at least one computational test item consisting of the original test items, and at least one split test item consisting of the original test items, wherein the dataset includes at least one wafer identifier; Based on at least one wafer identifier, at least one original test item, at least one computational test item consisting of the original test items, and at least one split test item consisting of the original test items, determine multiple data processing tasks and the dependencies between the data processing tasks. Based on the dependencies between the data processing tasks and the dataset, each data processing task is executed sequentially to obtain the data processing results corresponding to each data processing task. Based on at least one wafer identifier, at least one original test item, at least one computational test item consisting of the original test items, and at least one splitting test item consisting of the original test items, determine multiple data processing tasks and the dependencies between the data processing tasks, including: Based on each computational test item and at least one wafer identifier, multiple first-type data processing tasks are generated, wherein the total number of first-type data processing tasks is the same as the number of computational test items; Based on each split test item and at least one wafer identifier, multiple second-type data processing tasks are generated, wherein the total number of second-type data processing tasks is the same as the number of split test items, and the first-type data processing tasks are the precursor tasks of the second-type data processing tasks.

2. The method according to claim 1, characterized in that, Based on the dependencies between the data processing tasks and the dataset, each data processing task is executed sequentially to obtain the data processing results corresponding to each data processing task, including: Each data processing task is sliced ​​according to the total number of wafer identifiers to obtain multiple data processing task slices and the dependencies between each data processing task slice, wherein the total number of data processing task slices is the same as the total number of wafer identifiers. Based on the dependencies of each data processing task slice and the dataset, each data processing task slice is executed sequentially to obtain the data processing results corresponding to each data processing task slice. Based on the data processing results corresponding to each data processing task slice and the dependencies between each data processing task slice, the data processing results corresponding to each data processing task are determined.

3. The method according to claim 2, characterized in that, Data processing tasks include: Type I data processing tasks and Type II data processing tasks; Based on the total number of wafer identifiers, each data processing task is sliced ​​to obtain multiple data processing task slices and the dependencies between each data processing task slice, including: Based on the total number of wafer identifiers, each first type of data processing task is sliced ​​to obtain multiple data processing task slices corresponding to the first type of data processing task. Based on the total number of wafer identifiers, each second type of data processing task is sliced ​​to obtain multiple data processing task slices corresponding to the second type of data processing task. Based on the dependency relationship between the first type of data processing task and the second type of data processing task, determine the dependency relationship between multiple data processing task slices corresponding to the first type of data processing task and multiple data processing task slices corresponding to the second type of data processing task.

4. The method according to claim 1, characterized in that, Obtain the dataset, including: The main interface is displayed in response to a preset trigger action. In response to a trigger action on the New Dataset control in the main interface, display the dataset selection interface; In response to a trigger operation on the preview control in the dataset selection interface, the dataset corresponding to the test attribute information input by the user is obtained and the dataset is displayed in the dataset selection interface. The test attribute information includes at least one of the following: test stage, test time, batch number, test batch, and wafer identifier.

5. The method according to claim 4, characterized in that, Obtain at least one raw test item, including: In response to a trigger operation on the next control in the dataset selection interface, a test item selection interface is displayed, wherein the test item selection interface includes a query area and a test item list display area; In response to a trigger operation on a control in the query area, the list of test items in the test item list display area is updated according to the query information entered by the user; In response to a selection operation for a test item in the test item list, the selected test item is identified as the original test item.

6. The method according to claim 5, characterized in that, Obtain at least one computational test item consisting of the original test items, including: In response to the triggering operation of the next step control in the test item selection interface, the test item setting interface is displayed; In response to a triggering operation of adding a new calculation test item control in the calculation test item settings interface, a calculation test item editing interface is displayed, wherein the calculation test item editing interface includes: a first editing area, a first test item list display area, and a first function display area; In response to a trigger operation on a test item in the first test item list display area, the triggered test item is displayed in the first editing area; in response to a trigger operation on a function in the first function display area, the triggered function is displayed in the first editing area. In response to a trigger operation on the confirmation control in the calculation test item editing interface, a calculation test item is generated based on the triggered test item and the triggered function displayed in the first editing area.

7. The method according to claim 6, characterized in that, Obtain at least one split test item consisting of the original test item, including: In response to the triggering operation of the next step control in the calculation test item settings interface, the split test item settings interface is displayed; In response to the triggering operation of adding a new split test item control in the split test item settings interface, a split test item editing interface is displayed, wherein the split test item editing interface includes: a second editing area, a second test item list display area, and a second function display area; In response to a trigger operation on a test item in the second test item list display area, the triggered test item is displayed in the second editing area; in response to a trigger operation on a function in the second function display area, the triggered function is displayed in the editing area. In response to a trigger operation on the confirm control in the split test item editing interface, a split test item is generated based on the triggered test item and the triggered function displayed in the second editing area.

8. The method according to claim 1, characterized in that, Based on the dependencies between the data processing tasks and the dataset, each data processing task is executed sequentially to obtain the data processing results corresponding to each data processing task, including: After performing the first type of data processing task, perform the second type of data processing task to obtain a list of processing results; The execution of the first type of data processing task includes: obtaining each original test item contained in the calculation test item corresponding to the first type of data processing task; querying the dataset to obtain the test results of each original test item corresponding to the wafer identifier; performing a first calculation on the test results of each original test item corresponding to the wafer identifier according to the function in the calculation test item to obtain the processing result; and generating a first list based on the processing result. Executing the second type of data processing task includes: obtaining each original test item contained in the split test item corresponding to the second type of data processing task; performing a second calculation on the identification information of each original test item according to the function in the split test item to obtain attribute information; adding the attribute information to the first list to obtain a processing result list.

9. The method according to claim 1, characterized in that, Based on the dependencies between the data processing tasks and the dataset, each data processing task is executed sequentially to obtain the data processing results corresponding to each data processing task, including: If all the preceding tasks of the current data processing task have been completed, then the current data processing task is executed based on the dataset to obtain the data processing result corresponding to the current data processing task. The preceding tasks are the data processing tasks that the current data processing task depends on.

10. A data processing apparatus, characterized in that, include: An acquisition module is used to acquire a dataset, at least one original test item, at least one computational test item consisting of the original test items, and at least one split test item consisting of the original test items, wherein the dataset includes at least one wafer identifier; The task generation module is used to determine multiple data processing tasks and the dependencies between them based on at least one wafer identifier, at least one original test item, at least one computational test item composed of the original test items, and at least one split test item composed of the original test items. The task execution module is used to execute each data processing task sequentially according to the dependencies between the data processing tasks and the dataset, and to obtain the data processing results corresponding to each data processing task. The task generation module is specifically used for: Based on each computational test item and at least one wafer identifier, multiple first-type data processing tasks are generated, wherein the total number of first-type data processing tasks is the same as the number of computational test items; Based on each split test item and at least one wafer identifier, multiple second-type data processing tasks are generated, wherein the total number of second-type data processing tasks is the same as the number of split test items, and the first-type data processing tasks are the precursor tasks of the second-type data processing tasks.

11. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that is executed by the at least one processor to enable the at least one processor to perform the data processing method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the data processing method according to any one of claims 1-9.

13. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the data processing method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Data processing method, device and equipment and computer readable storage medium

    CN108628675A

  • Processing method and system for wafer detection data and storage medium

    CN115269175A