Model training methods, classification methods, devices, electronic equipment, and media
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-13
- Publication Date
- 2026-08-14
AI Technical Summary
导致对空间物体的特性的分析结果不理想
[0013]根据本公开的第六方面,提供了一种存储有计算机指令的非瞬时计算机可读存储介质,该计算机指令用于使计算机执行根据本公开提供的方法。
Smart Images

Figure CN116245159B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of ground-based optoelectronic observation technology, and in particular to the fields of deep learning, artificial intelligence and other technologies. In particular, it relates to a training method, classification method, device, electronic device, storage medium and program product for a deep learning model. Background Technology
[0002] Existing photometric processing methods can analyze the characteristics of space objects. However, these methods are limited to processing photometric data. Therefore, deep learning models can only process single-dimensional photometric data, thus ignoring additional information in the observation data. This results in unsatisfactory analysis results of the characteristics of space objects. Summary of the Invention
[0003] This disclosure provides a method for training a deep learning model, a classification method, an apparatus, an electronic device, a storage medium, and a program product.
[0004] According to a first aspect of this disclosure, one embodiment provides a method for training a deep learning model. The deep learning model includes an encoding module and a feature clustering module. The method includes: acquiring sample input data, the sample input data including sample photometric sequence data of an observed spatial object and additional information sequence data of at least one dimension corresponding to the sample photometric sequence data; processing the sample photometric sequence data and the additional information sequence data of at least one dimension using the encoding module to obtain encoded combined sequence data; processing the combined sequence data using the feature clustering module to obtain a clustering vector; determining the clustering result corresponding to the sample input data based on the clustering vector; and training the deep learning model based on the difference between the clustering result and the true clustering result of the sample input data.
[0005] For example, in the training method of the deep learning model provided in one embodiment of this disclosure, the additional information includes at least one of time information, distance information, azimuth elevation angle information, right ascension and declination information, and solar phase angle information; and the sample photometric sequence data includes photometric data at multiple times; wherein the time interval between any two adjacent photometric data is the same.
[0006] For example, in a deep learning model training method provided in an embodiment of this disclosure, the encoding module includes a nonlinear transformation network, and the additional information sequence data of at least one dimension includes additional information sequence data of multiple dimensions. Processing the sample photometric sequence data and the additional information sequence data of at least one dimension using the encoding module to obtain encoded combined sequence data includes: processing the additional information sequence data of multiple dimensions using a nonlinear transformation network to obtain coefficient encoded data; and obtaining encoded combined sequence data based on the sample photometric sequence data and the coefficient encoded data.
[0007] For example, in a deep learning model training method provided in an embodiment of this disclosure, the feature clustering module includes a first feature extraction network and a second feature extraction network; processing combined sequence data using the feature clustering module to obtain a clustering vector includes: processing the combined sequence data using the first feature extraction network and the second feature extraction network respectively to obtain a first feature vector and a second feature vector; and obtaining a clustering vector based on the first feature vector and the second feature vector.
[0008] For example, in a deep learning model training method provided in an embodiment of this disclosure, the feature clustering module includes multiple clustering networks corresponding to multiple predetermined clustering results; obtaining a clustering vector based on a first feature vector and a second feature vector includes: obtaining a third feature vector based on the first feature vector and the second feature vector; inputting the third feature vector into each clustering network to obtain multiple predicted values corresponding to the multiple clustering networks; and using the multiple predicted values as the clustering vector.
[0009] According to a second aspect of this disclosure, this disclosure provides a classification method, the method comprising: inputting target input data into a deep learning model to obtain target clustering results, wherein the deep learning model is trained using the method provided in this disclosure.
[0010] According to a third aspect of this disclosure, one embodiment provides a training apparatus for a deep learning model. The apparatus includes: an acquisition module for acquiring sample input data, the sample input data including sample photometric sequence data of an observed spatial object and additional information sequence data of at least one dimension corresponding to the sample photometric sequence data; a first acquisition module for processing the sample photometric sequence data and the additional information sequence data of at least one dimension using an encoding module to obtain encoded combined sequence data; a second acquisition module for processing the combined sequence data using a feature clustering module to obtain a clustering vector; a determination module for determining a clustering result corresponding to the sample input data based on the clustering vector; and a training module for training a deep learning model based on the difference between the clustering result and the actual clustering result of the sample input data.
[0011] According to a fourth aspect of this disclosure, one embodiment of this disclosure provides a classification apparatus, the apparatus comprising: a third obtaining module, configured to input target input data into a deep learning model to obtain target clustering results; wherein the deep learning model is trained using the apparatus provided in this disclosure.
[0012] According to a fifth aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method provided according to this disclosure.
[0013] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods provided according to this disclosure.
[0014] According to a seventh aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method provided according to this disclosure.
[0015] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0016] Figure 1 This is a flowchart of a training method for a deep learning model according to an embodiment of the present disclosure;
[0017] Figure 2 This is a schematic diagram of a deep learning model according to an embodiment of the present disclosure;
[0018] Figure 3 This is a schematic diagram of data processing according to an embodiment of the present disclosure;
[0019] Figure 4 This is a schematic diagram of data conversion according to an embodiment of the present disclosure;
[0020] Figure 5 This is a schematic diagram of data preprocessing according to an embodiment of the present disclosure;
[0021] Figure 6 This is a flowchart of a classification method according to an embodiment of the present disclosure;
[0022] Figure 7 This is a block diagram of a training apparatus for a deep learning model according to an embodiment of the present disclosure;
[0023] Figure 8 This is a block diagram of a classification apparatus according to an embodiment of the present disclosure; and
[0024] Figure 9 This is a block diagram of an electronic device according to an embodiment of the present disclosure, to which training methods and / or classification methods for deep learning models can be applied. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. Based on the described embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure. It should be noted that throughout the accompanying drawings, the same elements are represented by the same or similar reference numerals. In the following description, some specific embodiments are used for descriptive purposes only and should not be construed as limiting this disclosure in any way, but are merely examples of embodiments of this disclosure. Conventional structures or configurations will be omitted where they may cause confusion in understanding this disclosure. It should be noted that the shapes and dimensions of the components in the figures do not reflect actual size and proportion, but are only schematic representations of the embodiments of this disclosure.
[0026] Unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure shall have the ordinary meaning as understood by those skilled in the art. The terms "first," "second," and similar words used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components.
[0027] Figure 1 This is a flowchart of a training method for a deep learning model according to an embodiment of the present disclosure.
[0028] like Figure 1 As shown, the method 100 may include operations S110 to S150.
[0029] In operation S110, sample input data is acquired, which includes sample photometric sequence data for the observed spatial object and additional information sequence data of at least one dimension corresponding to the sample photometric sequence data.
[0030] In operation S120, the encoding module processes the sample photometric sequence data and at least one dimension of additional information sequence data to obtain the encoded combined sequence data.
[0031] In operation S130, the feature clustering module is used to process the combined sequence data to obtain cluster vectors.
[0032] In operation S140, the clustering result corresponding to the sample input data is determined based on the clustering vector.
[0033] In operation S150, a deep learning model is trained based on the difference between the clustering results and the actual clustering results of the sample input data.
[0034] Space objects refer to man-made objects orbiting the Earth in near-Earth space, typically including satellites, rocket debris, and other fragments. Observing space objects using photoelectric telescopes is one method to obtain measurement information such as their position and luminosity. The luminosity of a space object is the brightness of the sunlight it reflects as measured in an image acquired by a photoelectric telescope. Generally, photoelectric telescope measurements of space objects are performed continuously over a period of time, obtaining a series of time-series luminosity measurements. Simultaneously, a series of additional time-series information can be recorded, such as time, distance, azimuth, altitude, right ascension, declination, solar phase angle, etc. The sequentially recorded luminosity data of the space object constitutes the luminosity sequence data, and the sequentially recorded additional information at corresponding moments constitutes the additional information sequence data. This data can be arranged chronologically, recording multidimensional data composed of luminosity measurements and additional information at a given moment. By processing the luminosity sequence data of the space object, certain characteristic information of the observed space object can be analyzed, such as its shape and surface material characteristics.
[0035] The deep learning model in this embodiment includes an encoding module and a feature clustering module. The encoding module can process large-scale sample input data, which can be multi-dimensional sequence data. For example, it can include sample photometric sequence data of the observed spatial object and additional information sequence data of at least one dimension corresponding to the sample photometric sequence data. The feature clustering module can cluster spatial objects according to their pose selection features and typical external shape features.
[0036] Figure 2 This is a schematic diagram of a deep learning model according to an embodiment of the present disclosure. See also: Figure 2 The deep learning model 220 includes an encoding module 221 and a feature clustering module 222. For example, during model training, sample input data 210 can be input into the deep learning model. This sample input data 210 includes sample photometric sequence data for the observed spatial object and additional information sequence data corresponding to at least one dimension of the sample photometric sequence data. The additional information sequence data can be any additional information sequence data corresponding to the sample photometric sequence data, such as time sequence data, distance sequence data, azimuth and altitude angle sequence data, right ascension and declination sequence data, and solar phase angle sequence data, etc. The sample input data 210 is input to the encoding module 221, which outputs encoded combined sequence data. This combined sequence data is then input to the feature clustering module 220, which outputs a clustering result 230. The deep learning model is then trained based on the difference between the clustering result 230 and the true clustering result 250 (the labels of the sample input data) to obtain a trained deep learning model.
[0037] The deep learning model of this disclosure will be further described below with reference to relevant embodiments.
[0038] In some embodiments, the encoding module includes a nonlinear transformation network, and the additional information sequence data of at least one dimension includes additional information sequence data of multiple dimensions. Processing the sample photometric sequence data and the additional information sequence data of at least one dimension using the encoding module to obtain encoded combined sequence data includes: processing the additional information sequence data of multiple dimensions using a nonlinear transformation network to obtain coefficient encoded data; and obtaining encoded combined sequence data based on the sample photometric sequence data and the coefficient encoded data.
[0039] Figure 3 This is a schematic diagram of data processing according to an embodiment of the present disclosure, see [link / reference]. Figure 3 .
[0040] The sample photometric sequence data 320 and the additional information sequence data 310 of multiple dimensions can be arranged at equal intervals in chronological order. The additional information sequence data 310 of multiple dimensions is transformed into coefficient encoded data 330 by a deep learning-based nonlinear transformation network in the encoding module. The coefficient encoded data 330 and the sample photometric sequence data 320 can be processed using an element-wise multiplication operator in the encoding module to obtain the encoded combined sequence data 340, such as multiplying the coefficient encoded data 330 and the sample photometric sequence data 320 to generate the encoded combined sequence data 340. The encoded combined sequence data 340 is then processed by a feature clustering module to obtain a clustering vector 350; for example, the combined sequence data can be processed using a first feature extraction network and a second feature extraction network to obtain a first feature vector and a second feature vector; and the clustering vector is obtained based on the first feature vector and the second feature vector.
[0041] Figure 4 This is a schematic diagram illustrating data conversion according to an embodiment of the present disclosure. See also: Figure 4 .
[0042] There are d-dimensional additional information sequence data 410, with N additional information sequence data for each dimension. The d-dimensional additional information sequence data 410 is processed by a nonlinear transformation network 420 to obtain N*1 dimensional coefficient encoded data 430. For example, the nonlinear transformation network 420 can be a fully connected neural network with hidden layers. The nonlinear activation function at the network nodes gives the nonlinear transformation network 420 nonlinear transformation capability, thus transforming the fixed-length N-dimensional d-dimensional additional information sequence data 410 into N*1 dimensional sequence data of the same fixed length. The nonlinear activation function can be the tanh function or a function such as ReLU.
[0043] The coefficient-encoded data and sample photometric sequence data are further processed by the element-wise multiplication operator to synthesize the encoded combined sequence data, which is also an N*1 dimensional sequence data with a fixed length of N.
[0044] It can be understood that multiple elements in a clustering vector can represent the probabilities of different clustering results. The clustering result corresponding to the maximum value of a component in the clustering vector can be used as the clustering result corresponding to the sample input data. This allows the deep learning model to obtain clustering results, which can include classes such as spinning rocket debris, satellites with spinning box-wing shape features, satellites with stable attitude box-wing shape features, space debris with uncertain attitude and shape, and other classes.
[0045] In some embodiments, the feature clustering module includes a first feature extraction network and a second feature extraction network; processing combined sequence data using the feature clustering module to obtain a clustering vector includes: processing the combined sequence data using the first feature extraction network and the second feature extraction network respectively to obtain a first feature vector and a second feature vector; and obtaining a clustering vector based on the first feature vector and the second feature vector.
[0046] In some embodiments, the feature clustering module includes multiple clustering networks corresponding to multiple predetermined clustering results; obtaining a clustering vector based on a first feature vector and a second feature vector includes: obtaining a third feature vector based on the first feature vector and the second feature vector; inputting the third feature vector into each clustering network to obtain multiple predicted values corresponding to the multiple clustering networks; and using the multiple predicted values as the clustering vector.
[0047] For example, the first feature extraction network can be a temporal feature extraction network based on deep learning, and the second feature extraction network can be a spatial feature extraction network based on deep learning.
[0048] For example, the first feature extraction network can be constructed based on a one-dimensional feature extraction neural network to provide temporal sequence features of the data. The first feature extraction network is then used to process the combined sequence data to obtain a first feature vector, such as a temporal feature vector.
[0049] For example, the second feature extraction network can be constructed based on a two-dimensional feature extraction neural network to provide spatial autocorrelation features of the data. This second feature extraction network can be used to process combined sequence data to obtain a second feature vector, such as a spatial feature vector.
[0050] For example, a third feature vector can be obtained by concatenating the first and second feature vectors.
[0051] The first feature extraction network can be composed of several convolution operators and pooling operators linked together, ultimately obtaining the first half of the third feature vector, such as the first four dimensions.
[0052] The second feature extraction network can be composed of an initial rearrangement operator, several convolution operators, pooling subnetting, and an end rearrangement operator, ultimately yielding the latter half of the third feature vector, such as the last four dimensions.
[0053] Furthermore, the third feature vector is input into each clustering network to obtain multiple predicted values corresponding to the multiple clustering networks. The single-layer clustering networks can be fully connected neural networks with hidden layers, and non-linear activation functions can be used at the network nodes. Each single-layer clustering network outputs a specific value of the clustering vector, and the clustering vector is obtained by combining the multiple output values from the multiple clustering networks.
[0054] For cluster vectors, softmax can be used for normalization, as shown in Formula 1:
[0055]
[0056] The ordinal number of the largest component in the normalization result corresponds to the specific cluster number, which can be determined by the actual clustering results corresponding to the sample input data of the deep learning model. For example, the sample photometric sequence data used as sample input data includes cluster labels such as spinning rocket debris, spinning box-wing shaped satellites, stable-attitude box-wing shaped satellites, space debris with uncertain attitude and shape, and other categories. For example, Y1 corresponds to the spinning rocket debris category, Y2 corresponds to the spinning box-wing shaped satellite category, Y3 corresponds to the stable-attitude box-wing shaped satellite category, Y4 corresponds to the space debris with uncertain attitude and shape, and Y5 corresponds to other categories.
[0057] In some embodiments, the additional information includes at least one of time information, distance information, azimuth and altitude information, right ascension and declination information, and solar phase information; and the sample photometric sequence data includes photometric data at multiple times; wherein the time interval between any two adjacent photometric data is the same.
[0058] Figure 5 This is a schematic diagram of data preprocessing according to an embodiment of the present disclosure; that is, the sample input data is obtained after preprocessing the original data. The original data is typically a sequence with unequal intervals and variable time lengths. For large-scale spatial target photometric sequence data processing, in order to improve the efficiency of vectorized parallel computing, it is necessary to pre-encode the input data to ensure the uniformity of the time intervals and lengths of the sequence data. See also... Figure 5The original data 510 consists of photometric measurement observation results of space targets, such as time values (t1 to tn), photometric values (m1 to mn), and other additional data values (x1 to xn), which can be distance, azimuth and altitude angles, right ascension and declination, solar phase angle, etc., or default values. The time values (t1 to tn) may be unequally spaced, and for different space objects, their time lengths (i.e., the time difference between tn and t1) are also not fixed. A preprocessing unification method can be used to obtain equally spaced sequence data, such as sample photometric sequence data 521 and additional information sequence data 522 corresponding to at least one dimension of the sample photometric sequence data. For example, a fixed-length sampling sequence (e.g., with a time length of 255 seconds and an equal time interval (e.g., a 1-second step) is selected to determine a fixed-length sampling sequence (e.g., with a time length of 255 seconds and an equal time interval of 1 second, the fixed-length sampling sequence has N = 256 elements). For the raw data within the time interval between tn and t1, sampling is performed at equal time intervals (e.g., 1-second steps). For each sampling point Tx (x∈1,…,N), if data points exist (within half the step before and after the sampling point), their mean or the result of other filtering algorithms is used as the value of that sampling point Tx (Mx, Xx, etc.); if no data points exist, Tx is set to 0, and the value of that sampling point Tx (Mx, Xx, etc.) is set to zero. For raw data with insufficient time length (e.g., the time difference between tn and t1 is less than 255 seconds), Tx, Mx, Xx, etc., can be set to zero to supplement the sample photometric sequence data 521 and the additional information sequence data 522 corresponding to at least one dimension of the sample photometric sequence data; for data with a time length exceeding the sampling sequence length (e.g., the time difference between tn and t1 is greater than 255 seconds), the raw data can be truncated and discarded. After preprocessing, the sample input data is transformed into a fixed-length sequence, such as N=256. The sample photometric sequence data is a 256×1-dimensional sequence, and the additional information sequence data is a 256×d-dimensional sequence, where d is the dimension of the additional data and d≥1.
[0059] The network structure of deep learning models, including encoding and feature clustering modules, is scalable to data sequences of arbitrary length, observation frequency, and combinations. It adapts to sequences of arbitrary length and observation frequency by adjusting the fixed sampling time length and sampling interval of the input data. It also accommodates arbitrary combinations of data sequences by adjusting the dimensionality of the additional information sequence data. Deep learning models can further extend clustering results, such as by increasing the number of layers in the clustering network, thereby outputting more predictions and providing more comprehensive clustering predictions.
[0060] Figure 6 This is a flowchart of a classification method according to another embodiment of the present disclosure.
[0061] like Figure 6As shown, the method 600 may include operation S610.
[0062] When operating the S610, the target input data is fed into the deep learning model to obtain the target clustering results.
[0063] In embodiments of this disclosure, the deep learning model may be trained using the deep learning model training method provided in this disclosure. For example, the deep learning model may be trained using method 100.
[0064] In this embodiment of the disclosure, the target input data may be sequence data obtained from observing various space objects.
[0065] In this embodiment of the disclosure, the target clustering result can indicate the clustering result of spatial objects according to their attitude rotation characteristics and typical external shape characteristics.
[0066] Figure 7 This is a block diagram of a training apparatus for a deep learning model according to an embodiment of the present disclosure.
[0067] like Figure 7 As shown, the device 700 may include an acquisition module 710, a first acquisition module 720, a second acquisition module 730, a determination module 740, and a training module 750.
[0068] The module 710 is used to acquire sample input data, which includes sample photometric sequence data of the observed spatial object and additional information sequence data of at least one dimension corresponding to the sample photometric sequence data; the first acquisition module 720 is used to process the sample photometric sequence data and the additional information sequence data of at least one dimension using the encoding module to obtain encoded combined sequence data; the second acquisition module 730 is used to process the combined sequence data using the feature clustering module to obtain clustering vectors; the determination module 740 is used to determine the clustering result corresponding to the sample input data based on the clustering vectors; and the training module 750 is used to train a deep learning model based on the difference between the clustering result and the true clustering result of the sample input data.
[0069] In some embodiments, the additional information includes at least one of time information, distance information, azimuth and altitude information, right ascension and declination information, and solar phase information; and the sample photometric sequence data includes photometric data at multiple times; wherein the time interval between any two adjacent photometric data is the same.
[0070] In some embodiments, the encoding module includes a nonlinear transformation network, and the additional information sequence data of at least one dimension includes additional information sequence data of multiple dimensions. The first obtaining module includes: a first obtaining submodule, used to process the additional information sequence data of multiple dimensions using the nonlinear transformation network to obtain coefficient encoded data; and a second obtaining submodule, used to obtain encoded combined sequence data based on the sample photometric sequence data and the coefficient encoded data.
[0071] In some embodiments, the feature clustering module includes a first feature extraction network and a second feature extraction network; the second obtaining module includes: a third obtaining submodule, used to process combined sequence data using the first feature extraction network and the second feature extraction network respectively to obtain a first feature vector and a second feature vector; and a fourth obtaining submodule, used to obtain a clustering vector based on the first feature vector and the second feature vector.
[0072] In some embodiments, the feature clustering module includes multiple clustering networks corresponding to multiple predetermined clustering results; the fourth obtaining submodule includes: a first obtaining unit, used to obtain a third feature vector based on a first feature vector and a second feature vector; a second obtaining unit, used to input the third feature vector into each clustering network respectively to obtain multiple predicted values corresponding to the multiple clustering networks; and a third obtaining unit, used to use the multiple predicted values as clustering vectors.
[0073] Figure 8 This is a block diagram of a sorting apparatus according to an embodiment of the present disclosure.
[0074] like Figure 8 As shown, the device 800 may include a third acquisition module 810.
[0075] The third acquisition module is used to input the target input data into the deep learning model to obtain the target clustering results.
[0076] For example, a deep learning model is trained using the apparatus provided in this disclosure.
[0077] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0078] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0079] Figure 9A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0080] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0081] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0082] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as training methods and / or classification methods for deep learning models. For example, in some embodiments, the training methods and / or classification methods for deep learning models can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the training methods and / or classification methods for deep learning models described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured by any other suitable means (e.g., by means of firmware) to perform training and / or classification methods for deep learning models.
[0083] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0084] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0085] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0086] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) monitor or an LCD (liquid crystal display)) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0087] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0088] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0089] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0090] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for training a deep learning model, the deep learning model comprising an encoding module and a feature clustering module, the method comprising: Acquire sample input data, which includes sample photometric sequence data for the observed space object and additional information sequence data for at least one dimension corresponding to the sample photometric sequence data. The photometric of the observed space object is the brightness of the sunlight reflected by the observed space object in the image acquired by the photoelectric telescope. The sample photometric sequence data includes photometric data at multiple times, and the time interval between any two adjacent photometric data is the same. The encoding module is used to process the sample photometric sequence data and the additional information sequence data of at least one dimension to obtain encoded combined sequence data; The combined sequence data is processed using the feature clustering module to obtain clustering vectors; Based on the clustering vector, determine the clustering result corresponding to the sample input data; and The deep learning model is trained based on the difference between the clustering results and the actual clustering results of the sample input data. The encoding module includes a nonlinear transformation network, and the at least one-dimensional supplementary information sequence data includes supplementary information sequence data of multiple dimensions. The process of using the encoding module to process the sample photometric sequence data and the at least one-dimensional supplementary information sequence data to obtain the encoded combined sequence data includes: The nonlinear transformation network is used to process the additional information sequence data of the multiple dimensions to obtain coefficient encoded data; The coded coefficient data is multiplied by the sample photometric sequence data to obtain the coded combined sequence data.
2. The method according to claim 1, wherein, Additional information includes at least one of the following: time information, distance information, azimuth and altitude information, right ascension and declination information, and solar phase angle information.
3. The method according to claim 1, wherein, The feature clustering module includes a first feature extraction network and a second feature extraction network; the process of using the feature clustering module to process the combined sequence data to obtain clustering vectors includes: The combined sequence data is processed using the first feature extraction network and the second feature extraction network, respectively, to obtain a first feature vector and a second feature vector; and The clustering vector is obtained based on the first feature vector and the second feature vector.
4. The method according to claim 3, wherein, The feature clustering module includes multiple clustering networks corresponding to multiple predetermined clustering results; obtaining the clustering vector based on the first feature vector and the second feature vector includes: Based on the first feature vector and the second feature vector, the third feature vector is obtained; The third feature vector is input into each clustering network to obtain multiple predicted values corresponding to the multiple clustering networks; and The multiple predicted values are used as the clustering vector.
5. A classification method, comprising: Input the target data into the deep learning model to obtain the target clustering results; The deep learning model is trained using the method described in any one of claims 1 to 4.
6. A training apparatus for a deep learning model, the deep learning model comprising an encoding module and a feature clustering module, the apparatus comprising: The acquisition module is used to acquire sample input data, which includes sample photometric sequence data for the observed space object and additional information sequence data of at least one dimension corresponding to the sample photometric sequence data. The photometric of the observed space object is the brightness of the sunlight reflected by the observed space object in the image acquired by the photoelectric telescope. The sample photometric sequence data includes photometric data at multiple times, and the time interval between any two adjacent photometric data is the same. The first obtaining module is used to process the sample photometric sequence data and the additional information sequence data of at least one dimension using the encoding module to obtain encoded combined sequence data; The second obtaining module is used to process the combined sequence data using the feature clustering module to obtain a clustering vector; The determining module is configured to determine the clustering result corresponding to the sample input data based on the clustering vector; and The training module is used to train the deep learning model based on the difference between the clustering results and the actual clustering results of the sample input data; The encoding module includes a nonlinear transformation network, and the additional information sequence data of at least one dimension includes additional information sequence data of multiple dimensions. The first obtaining module includes: The first obtaining submodule is used to process the additional information sequence data of the multiple dimensions using the nonlinear transformation network to obtain coefficient encoded data; The second obtaining submodule is used to multiply the coefficient encoded data with the sample photometric sequence data to obtain the encoded combined sequence data.
7. A sorting device, comprising: The third acquisition module is used to input the target input data into the deep learning model to obtain the target clustering results; The deep learning model is trained using the apparatus described in claim 6.
8. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 5.
9. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Semi-supervised learning pseudo label assignment method based on clustering fusion
CN112418331A
Network encryption traffic classification method and system based on multi-feature learning
CN113037730A
Solar radiation prediction method and system based on double-branch feature extraction
CN115099461A