Classification method and apparatus for target objects
By acquiring Raman images of target objects, performing multi-scale feature extraction, and combining statistical probability matrix classification, the problem of low image recognition accuracy in medical scenarios is solved, achieving high accuracy and reliability in object state classification.
Patent Information
- Application Number
- CN202211726319.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2042-12-30
AI Technical Summary
In medical settings, existing image recognition technologies are not very accurate in recognizing medical images, resulting in low reliability in determining the state of objects.
By acquiring the first and second Raman images of the target object, multi-scale feature extraction and statistical probability matrix are used, combined with historical physiological data, to classify the target object's state.
It improves the accuracy and reliability of determining the state of objects, and significantly enhances the accuracy of classification by combining images and statistical probability matrices for classification.
Smart Images

Figure CN116229145B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method and apparatus for classifying target objects. Background Technology
[0002] With the development of computer technology, artificial intelligence technology is booming. As a branch of artificial intelligence, image recognition technology is being applied in an increasingly wide range of applications. For example, image recognition technology can be applied to facial recognition scenarios, where images containing faces can be identified to obtain identity information corresponding to the faces. Alternatively, it can be applied to medical scenarios, where medical images can be identified to discover lesions that are difficult for the human eye to distinguish, thus revealing the condition of the object and assisting doctors in determining treatment plans.
[0003] In related technologies, in medical scenarios, the accuracy of identifying medical images using only image recognition technology is not high, resulting in low reliability of the determined object's state. Summary of the Invention
[0004] This application provides a method and apparatus for classifying target objects, which can improve the reliability of the determined state of the object. The technical solution is as follows.
[0005] On the one hand, a method for classifying target objects is provided, the method comprising:
[0006] A first Raman image, a second Raman image, and a statistical probability matrix of a target object are acquired. The statistical probability matrix is used to reflect the probability that the target object is in multiple candidate states. The statistical probability matrix is determined based on the historical physiological data of the target object. The historical physiological data is used to record the preset states that the target object has been in. The first Raman image and the second Raman image have different resolutions.
[0007] Multi-scale feature extraction is performed on the first Raman image and the second Raman image to obtain multi-scale Raman features;
[0008] Based on the multi-scale Raman features, a first classification result for the target object is obtained;
[0009] Based on the first classification result and the statistical probability matrix, a second classification result of the target object is determined. The second classification result is the target state of the target object, and the target state belongs to the multiple candidate states.
[0010] In one possible implementation, acquiring the first Raman image, the second Raman image, and the statistical probability matrix of the target object includes:
[0011] Acquire Raman data of the target object, the Raman data including Raman signal values corresponding to multiple time points; increase the dimensionality of the Raman data to obtain the first Raman image and the second Raman image;
[0012] Obtain historical physiological data of the target object; based on the historical physiological data, determine a preset state that the target object was previously in; determine the statistical probability matrix corresponding to the preset state.
[0013] In one possible implementation, the step of upscaling the Raman data to obtain the first Raman image and the second Raman image includes:
[0014] The Raman data is downsampled using a first sampling rate and a second sampling rate to obtain first downsampled Raman data corresponding to the first sampling rate and second downsampled Raman data corresponding to the second sampling rate, wherein the first sampling rate and the second sampling rate are different.
[0015] The first downsampled Raman data and the second downsampled Raman data are transformed into polar coordinates to obtain the radius and angle of each data point in the first downsampled Raman data and the second downsampled Raman data in polar coordinates.
[0016] Based on the radius and angle of each data point in the first downsampled Raman data and the second downsampled Raman data in the polar coordinate system, the correlation between every two data points in the first downsampled Raman data and the second downsampled Raman data is determined;
[0017] The first Raman image and the second Raman image are generated based on the correlation between every two data points in the first downsampled Raman data and the second downsampled Raman data.
[0018] In one possible implementation, determining the statistical probability matrix corresponding to the preset state includes:
[0019] Obtain state statistics information, which includes the probability of being in the multiple candidate states given that the person has been in different preset states in the past.
[0020] Obtain the statistical probability matrix corresponding to the preset state from the state statistics information.
[0021] In one possible implementation, the multi-scale feature extraction of the first Raman image and the second Raman image to obtain multi-scale Raman features includes:
[0022] The first Raman image is convolved using a first convolution kernel to obtain the first Raman image features;
[0023] The second Raman image is convolved using a second convolution kernel to obtain the features of the second Raman image. The first and second convolution kernels have different sizes.
[0024] The first Raman image features and the second Raman image features are fused to obtain the multi-scale Raman features.
[0025] In one possible implementation, the classification based on the multi-scale Raman features to obtain a first classification result for the target object includes:
[0026] A fully connected layer is applied to the multi-scale Raman features to obtain a first classification result for the target object, which includes the probability that the target object is in multiple candidate states.
[0027] In one possible implementation, determining the second classification result of the target object based on the first classification result and the statistical probability matrix includes:
[0028] The first classification result and the statistical probability matrix are concatenated to obtain a multimodal matrix;
[0029] The multimodal matrix is multiplied by the weight matrix to obtain the prediction matrix, wherein the multimodal matrix and the weight matrix have the same size.
[0030] Based on the prediction matrix, the second classification result of the target object is determined.
[0031] In one possible implementation, determining the second classification result of the target object based on the prediction matrix includes:
[0032] The prediction matrix is summed according to the dimensions of the preset states to obtain the classification matrix;
[0033] The classification matrix is normalized to obtain the state probability distribution column of the target object, which includes the probability that the target object is in multiple candidate states;
[0034] Based on the state probability distribution, the second classification result of the target object is determined.
[0035] In one possible implementation, determining the second classification result of the target object based on the state probability distribution column includes:
[0036] The candidate state with the highest probability in the state probability distribution column is determined as the second classification result of the target object.
[0037] On the one hand, a classification device for target objects is provided, the device comprising:
[0038] The acquisition module is used to acquire a first Raman image, a second Raman image, and a statistical probability matrix of a target object. The statistical probability matrix is used to reflect the probability that the target object is in multiple candidate states. The statistical probability matrix is determined based on the historical physiological data of the target object. The historical physiological data is used to record the preset states that the target object has been in. The first Raman image and the second Raman image have different resolutions.
[0039] The feature extraction module is used to perform multi-scale feature extraction on the first Raman image and the second Raman image to obtain multi-scale Raman features;
[0040] The first classification module is used to classify based on the multi-scale Raman features to obtain the first classification result of the target object;
[0041] The second classification module is used to determine the second classification result of the target object based on the first classification result and the statistical probability matrix. The second classification result is the target state of the target object, and the target state belongs to the multiple candidate states.
[0042] In one possible implementation, the acquisition module is configured to acquire Raman data of the target object, the Raman data including Raman signal values corresponding to multiple time points; perform dimensionality upscaling on the Raman data to obtain a first Raman image and a second Raman image; acquire historical physiological data of the target object; determine a preset state that the target object was previously in based on the historical physiological data; and determine the statistical probability matrix corresponding to the preset state.
[0043] In one possible implementation, the acquisition module is configured to downsample the Raman data using a first sampling rate and a second sampling rate to obtain first downsampled Raman data corresponding to the first sampling rate and second downsampled Raman data corresponding to the second sampling rate, wherein the first sampling rate and the second sampling rate are different; convert the first downsampled Raman data and the second downsampled Raman data to a polar coordinate system to obtain the radius and angle of each data point in the first downsampled Raman data and the second downsampled Raman data in the polar coordinate system; determine the correlation between every two data points in the first downsampled Raman data and the second downsampled Raman data based on the radius and angle of each data point in the first downsampled Raman data and the second downsampled Raman data in the polar coordinate system; and generate a first Raman image and a second Raman image based on the correlation between every two data points in the first downsampled Raman data and the second downsampled Raman data.
[0044] In one possible implementation, the acquisition module is used to acquire state statistics, which include the probability of being in the multiple candidate states given that the state has been in different preset states in the past; and to acquire the statistical probability matrix corresponding to the preset state from the state statistics.
[0045] In one possible implementation, the feature extraction module is used to perform convolution processing on the first Raman image using a first convolution kernel to obtain a first Raman image feature; to perform convolution processing on the second Raman image using a second convolution kernel to obtain a second Raman image feature, wherein the first convolution kernel and the second convolution kernel have different sizes; and to fuse the first Raman image feature and the second Raman image feature to obtain the multi-scale Raman feature.
[0046] In one possible implementation, the first classification module is used to perform a fully connected operation on the multi-scale Raman features to obtain a first classification result of the target object, wherein the first classification result includes the possibility that the target object is in multiple candidate states.
[0047] In one possible implementation, the second classification module is used to concatenate the first classification result and the statistical probability matrix to obtain a multimodal matrix; multiply the multimodal matrix with the weight matrix to obtain a prediction matrix, wherein the multimodal matrix and the weight matrix have the same size; and determine the second classification result of the target object based on the prediction matrix.
[0048] In one possible implementation, the second classification module is used to sum the prediction matrix according to the dimensions of a preset state to obtain a classification matrix; normalize the classification matrix to obtain a state probability distribution column of the target object, the state probability distribution column including the probability that the target object is in multiple candidate states; and determine a second classification result of the target object based on the state probability distribution column.
[0049] In one possible implementation, the second classification module is used to determine the candidate state corresponding to the highest probability in the state probability distribution column as the second classification result of the target object.
[0050] On one hand, a computer device is provided, the computer device including one or more processors and one or more memories, the one or more memories storing at least one piece of program code, the program code being loaded and executed by the one or more processors to implement the operations performed by the classification method of the target object.
[0051] On one hand, a computer-readable storage medium is provided, wherein at least one piece of program code is stored in the computer-readable storage medium, the program code being loaded and executed by a processor to implement the operations performed by the classification method of the target object.
[0052] The technical solution provided in this application acquires a first Raman image, a second Raman image, and a statistical probability matrix of a target object. The statistical probability matrix represents the probability that the target object is in multiple candidate states. This matrix is determined based on the target object's historical physiological data, which includes preset states the target object has previously been in. The first and second Raman images have different resolutions. Multi-scale feature extraction is performed on the first and second Raman images to obtain multi-scale Raman features. Classification is performed based on these multi-scale Raman features to obtain a first classification result for the target object. Based on the first classification result and the statistical probability matrix, a second classification result is determined for the target object, which is the target state in which the target object is located. Combining images and the statistical probability matrix for target object classification results in high accuracy. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is a schematic diagram illustrating the implementation environment of a target object classification method provided in an embodiment of this application;
[0055] Figure 2 This is a flowchart of a target object classification method provided in an embodiment of this application;
[0056] Figure 3 This is a flowchart of another method for classifying target objects provided in an embodiment of this application;
[0057] Figure 4 This is a flowchart illustrating how to determine a first classification result, as provided in an embodiment of this application.
[0058] Figure 5 This is a flowchart illustrating how to determine a second classification result, provided in an embodiment of this application.
[0059] Figure 6 This is a flowchart of another target object classification method provided in the embodiments of this application;
[0060] Figure 7This is a schematic diagram of an adaptive weight provided in an embodiment of this application;
[0061] Figure 8 This is a schematic diagram of the structure of a target object classification device provided in an embodiment of this application;
[0062] Figure 9 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application.
[0063] Figure 10 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation
[0064] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0065] In order to illustrate the technical solutions provided in the embodiments of this application, the terms involved in the embodiments of this application will be introduced below.
[0066] Artificial intelligence (AI) is the theory, methods, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0067] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge sub-models to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.
[0068] Raman effect: Light scatters elastically and inelastically when it strikes a substance. Elastic scattering produces light with the same wavelength as the excitation light. Inelastic scattering produces light with both wavelengths longer and shorter than the excitation light, collectively known as the Raman effect. When a gas, liquid, or transparent sample is irradiated with monochromatic light whose wavelength is much smaller than the sample particle size, most of the light is transmitted in its original direction, while a small portion is scattered at different angles, producing scattered light. When observed perpendicularly, in addition to Rayleigh scattering with the same frequency as the original incident light, there is a series of symmetrically distributed, very weak Raman spectral lines that are shifted from the incident light frequency. This phenomenon is called the Raman effect. Since the number of Raman spectral lines, the magnitude of the shift, and the length of the spectral lines are directly related to the vibrational or rotational energy levels of the sample molecules, this effect is crucial.
[0069] Raman spectroscopy (RS) is considered a rapid optical detection tool that can acquire unique spectral features similar to fingerprints for component identification within minutes. It also has the advantages of being non-invasive, easy to operate, and highly sensitive, and has achieved substantial success in various research fields.
[0070] A polar coordinate system is a coordinate system in a plane consisting of a pole, a polar axis, and a polar radius. A point O is chosen on the plane, called the pole. A ray Ox is drawn from O, called the polar axis. A unit length is then chosen, and angles are usually defined as positive counterclockwise. Thus, the position of any point P on the plane can be determined by the length ρ of the line segment OP and the angle θ from Ox to OP. The ordered pair (ρ, θ) is called the polar coordinates of point P, denoted as P(ρ, θ); ρ is called the polar radius of point P, and θ is called the polar angle of point P.
[0071] Cardiovascular disease (CVD): Cardiovascular and cerebrovascular diseases are a collective term for diseases of the heart and brain blood vessels. They broadly refer to ischemic or hemorrhagic diseases of the heart, brain, and other tissues caused by conditions such as hyperlipidemia, high blood viscosity, atherosclerosis, and hypertension. Cardiovascular disease is one of the leading causes of death worldwide. Coronary artery disease (CAD) is the leading type of death among CVD patients. Many studies have shown that CAD patients often develop atrial fibrillation (AF) or acute myocardial infarction (AMI), which are different subtypes of CVD.
[0072] Normalization: Mapping sequences of values with different ranges to the interval (0, 1) to facilitate data processing. In some cases, normalized values can be directly expressed as probabilities.
[0073] Learning rate: Used to control the learning progress of the model. The learning rate guides the model in adjusting network weights using the gradient of the loss function during gradient descent. If the learning rate is too large, the loss function may directly skip the global optimum, resulting in excessive loss. If the learning rate is too small, the loss function changes very slowly, greatly increasing the convergence complexity of the network and making it easy to get trapped in local minima or saddle points.
[0074] Embedded coding, mathematically speaking, represents a correspondence, that is, mapping data in space X to space Y using a function F. This function F is injective, and the mapping result preserves the structure. An injective function means that the mapped data uniquely corresponds to the original data, and preserving the structure means that the order of the original data remains the same. For example, if there are data X1 and X2 before mapping, after mapping we get Y1 corresponding to X1 and Y2 corresponding to X2. If the original data X1 > X2, then correspondingly, the mapped data Y1 > Y2. For words, this means mapping words to another space to facilitate subsequent machine learning and processing.
[0075] Attention weights represent the importance of a piece of data during training or prediction. Importance indicates the magnitude of the influence of input data on output data. Data with high importance corresponds to a higher attention weight, while data with low importance corresponds to a lower attention weight. The importance of data varies in different scenarios, and training the model to assign attention weights is essentially the process of determining the importance of that data.
[0076] After introducing the terms used in the embodiments of this application, the implementation environment of the embodiments of this application will be described below.
[0077] Figure 1 This is a schematic diagram illustrating the implementation environment of a target object classification method provided in this application embodiment. See also... Figure 1 The implementation environment may include terminal 110 and server 140.
[0078] Terminal 110 is connected to server 140 via a wireless or wired network. Optionally, terminal 110 may be a smartphone, tablet, laptop, desktop computer, etc., but is not limited to these. Terminal 110 has an application installed and running that supports target object classification.
[0079] Server 140 is a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. Server 140 provides background services for applications running on the terminal. For example, server 140 can train an object classification model to classify target objects for use by terminal 110.
[0080] Those skilled in the art will understand that the number of terminals described above can be more or less. For example, there may be only one terminal, or there may be dozens or hundreds of terminals, or even more, in which case other terminals may also be included in the above implementation environment. This application does not limit the number of terminals or the type of device in its embodiments.
[0081] After introducing the implementation environment of the target object classification method provided in the embodiments of this application, the application scenarios of the target object classification method provided in the embodiments of this application will be described below. It should be noted that the terminal in the following description is the terminal 110 in the above implementation environment, and the server is the server 140 in the above implementation environment.
[0082] The target object classification method provided in this application embodiment can be applied to scenarios where multiple types of objects are classified, such as patients, vehicles, or buildings. This application embodiment does not limit this application.
[0083] When the target object classification method provided in this application is applied to a patient classification scenario, the terminal acquires a first Raman image, a second Raman image, and a statistical probability matrix of the target object, i.e., the patient. Both the first and second Raman images are obtained based on Raman data of the target object, which is obtained by performing Raman testing on a blood sample of the target object. The statistical probability matrix reflects the probability of the target object being in multiple candidate states, i.e., multiple candidate diseases. This statistical probability matrix is obtained based on the target object's historical physiological data, which records the target object's past preset states, i.e., the target object's medical history, where the preset state is a preset disease. Multi-scale feature extraction is performed on the first and second Raman images to obtain multi-scale Raman features. Classification is performed based on these multi-scale Raman features to obtain a first classification result for the target object, which includes the probability that the target object is in multiple candidate states, i.e., multiple candidate diseases. Based on the first classification result and the statistical probability matrix, a second classification result for the target object is determined, which is the target disease among the multiple candidate diseases.
[0084] It should be noted that the above description is based on the application of the target object classification method provided in the embodiments of this application to classify patients. In other scenarios, the implementation process of the method belongs to the same inventive concept as described above, and will not be repeated here.
[0085] After introducing the implementation environment and application scenarios of the embodiments of this application, the technical solutions provided by the embodiments of this application are described below. (See also...) Figure 2 Taking the terminal as the executing entity as an example, the method includes the following steps.
[0086] 201. The terminal acquires a first Raman image, a second Raman image, and a statistical probability matrix of the target object. The statistical probability matrix is used to reflect the probability that the target object is in multiple candidate states. The statistical probability matrix is determined based on the historical physiological data of the target object. The historical physiological data is used to record the preset states that the target object was previously in. The first Raman image and the second Raman image have different resolutions.
[0087] The target object is the object to be classified, such as a patient, vehicle, or building. In the following description, a patient will be used as an example. The first Raman image and the second Raman image have different resolutions; that is, one of the first and second Raman images has a higher resolution, and the other has a lower resolution. The statistical probability matrix is based on statistics and can reflect the correlation between being in multiple candidate states and other factors to a certain extent. In this embodiment, the statistical probability matrix is related to the historical physiological data of the target object, which records the preset states the target object has previously been in. Accordingly, the statistical probability matrix can reflect the correlation between the preset states the target object has previously been in and being in multiple candidate states.
[0088] 202. The terminal performs multi-scale feature extraction on the first Raman image and the second Raman image to obtain multi-scale Raman features.
[0089] Among them, multi-scale feature extraction can extract features from the first Raman image and the second Raman image in different dimensions, and the resulting multi-scale Raman features have stronger expressive power.
[0090] 203. The terminal performs classification based on the multi-scale Raman features to obtain the first classification result of the target object.
[0091] Among them, classification based on the multi-scale Raman features, that is, classification of target objects based on Raman images, the first classification result can also be called the Raman image classification result.
[0092] 204. Based on the first classification result and the statistical probability matrix, the terminal determines the second classification result of the target object. The second classification result is the target state of the target object, and the target state belongs to the multiple candidate states.
[0093] The second classification result is obtained based on the first classification result and the statistical probability matrix. The first classification result is the Raman image classification result, and the statistical probability matrix is equivalent to the statistical classification result. The second classification result integrates the two results and has higher accuracy.
[0094] The technical solution provided in this application acquires a first Raman image, a second Raman image, and a statistical probability matrix of a target object. The statistical probability matrix represents the probability that the target object is in multiple candidate states. This matrix is determined based on the target object's historical physiological data, which includes preset states the target object has previously been in. The first and second Raman images have different resolutions. Multi-scale feature extraction is performed on the first and second Raman images to obtain multi-scale Raman features. Classification is performed based on these multi-scale Raman features to obtain a first classification result for the target object. Based on the first classification result and the statistical probability matrix, a second classification result is determined for the target object, which is the target state in which the target object is located. Combining images and the statistical probability matrix for target object classification results in high accuracy.
[0095] Steps 201-204 above are a brief introduction to the technical solutions provided in the embodiments of this application. The technical solutions provided in the embodiments of this application will be explained more clearly below with some examples. See [link to relevant documentation]. Figure 3 Taking the server as the executing entity as an example, the method includes the following steps.
[0096] 301. The server obtains Raman data of the target object, which includes Raman signal values corresponding to multiple time points.
[0097] The target object is the object to be classified, such as a patient, vehicle, or building. In the following description, a patient as the target object will be used as an example. In this embodiment, the patient refers to a patient with cardiovascular disease (CVD). Classifying the patient involves determining the subtype of CVD, which includes coronary artery disease (CAD), atrial fibrillation (AF), and acute myocardial infarction (AMI), among others. A healthy control group (CON) is also included for comparison. In clinical diagnosis, these subtypes should be accurately classified as soon as possible to enable appropriate treatment and improve survival rates. Over the past decade, methods such as electrocardiography, cardiac enzyme profiles, cardiac computed tomography (CT), and coronary angiography have been widely used to diagnose these conditions. These diagnostic methods are time-consuming in data acquisition and analysis. Generally, diagnosis and treatment within the first hour or "golden hour" after a heart attack are crucial for patients, as this can save their lives.
[0098] Raman data is obtained by performing Raman testing on a blood sample of the target object; the Raman signal value is the signal value obtained during the Raman test. This Raman test is used to measure the Raman spectrum (RS) of the blood sample, and the data corresponding to this Raman spectrum is the Raman data. Raman spectroscopy is considered a rapid optical detection tool that can acquire unique spectral features similar to fingerprints for component identification within minutes. It also boasts advantages such as being non-invasive, simple to operate, and highly sensitive, achieving substantial success in various research fields.
[0099] In one possible implementation, the server obtains the Raman data of the target object from a terminal connected to a Raman testing device. The Raman testing device performs Raman testing on the blood sample of the target object, and the Raman data obtained is sent to the terminal. The terminal then uploads the Raman data to the server. The server then obtains the Raman data of the target object.
[0100] It should be noted that the terminal uploads the Raman data of the target object to the server, and the server subsequently processes the Raman data, with the consent of the target object.
[0101] In this implementation, the server can directly obtain Raman data from the terminal, resulting in high efficiency in acquiring Raman data.
[0102] In one possible implementation, the server retrieves the Raman data of the target object from a Raman database that stores Raman data of multiple objects.
[0103] In this implementation, the server can obtain Raman data of the target object from the Raman database, and the acquisition efficiency of Raman data is relatively high.
[0104] 302. The server performs dimensionality upscaling on the Raman data to obtain the first Raman image and the second Raman image, which have different resolutions.
[0105] In this context, Raman data is one-dimensional, specifically a one-dimensional time series. Upscaling this Raman data transforms it into two-dimensional data, which can be represented as images, namely the first Raman image and the second Raman image. The first and second Raman images have different resolutions; one has a higher resolution than the other. Images with different resolutions contain a variety of information. In computer vision, high-resolution images contain more detailed information, while low-resolution images contain more abstract knowledge. Furthermore, in Raman images, high resolution preserves more local cosine relationships between points, while low resolution expresses more global semantic connections between different groups. Since the first and second Raman images represent two-dimensional Raman data, their resolutions are related to the sampling rate of the two-dimensional Raman data. In some embodiments, the resolutions of the first and second Raman images are 32 and 64, respectively. In some embodiments, step 302 above can be implemented by the multi-scale feature extraction module of the object classification model.
[0106] In one possible implementation, the server downsamples the Raman data using a first sampling rate and a second sampling rate, obtaining first downsampled Raman data corresponding to the first sampling rate and second downsampled Raman data corresponding to the second sampling rate, wherein the first and second sampling rates are different. The server transforms the first and second downsampled Raman data into a polar coordinate system, obtaining the radius and angle of each data point in the first and second downsampled Raman data in the polar coordinate system. Based on the radius and angle of each data point in the first and second downsampled Raman data in the polar coordinate system, the server determines the correlation between every two data points in the first and second downsampled Raman data. Based on the correlation between every two data points in the first and second downsampled Raman data, the server generates the first Raman image and the second Raman image.
[0107] The first sampling rate and the second sampling rate are different. When sampling Raman data at different sampling rates, it means sampling at different intervals of data points in the Raman data. The more data points at intervals, the lower the sampling rate; the fewer data points at intervals, the higher the sampling rate.
[0108] In this implementation, the server first downsamples the Raman data using different sampling rates. The resulting first and second downsampled Raman data are then converted to polar coordinates, yielding the radius and angle of each data point in both sets of data in the polar coordinates. Using these radius and angle values, the correlation between any two data points is determined. Based on this correlation, the Raman data can be upsampled to obtain the first and second Raman images.
[0109] For example, the server standardizes the Raman data to obtain standardized target Raman data. The server then downsamples the target Raman data using a first sampling rate and a second sampling rate, obtaining first downsampled Raman data and second downsampled Raman data. For the first downsampled Raman data, the server transforms it to a polar coordinate system, obtaining the radius and angle of each data point in the first downsampled Raman data in the polar coordinate system. Based on the radius and angle of each data point in the first downsampled Raman data in the polar coordinate system, the server determines the correlation between every two data points in the first downsampled Raman data. Based on the correlation between every two data points in the first downsampled Raman data, the server generates the first Raman image. The method for generating the second Raman image belongs to the same inventive concept as the method for generating the first Raman image, and the implementation process will not be described in detail.
[0110] For example, servers use the Gram corner field method to increase the dimensionality of Raman data. Inspired by the successful use of deep learning in computer vision and natural language processing, some scholars have proposed a framework for encoding time series data as an image type, called Gram corner field (GAF). By using polar coordinates, GAF represents an image as a Gram matrix, where each element is the triangular sum between different sample points.
[0111] The server standardizes the Raman data, that is, converts the Raman data into a standardized one-dimensional spectral sequence, which can be achieved by the following formula (1). The standardized one-dimensional spectral sequence is the target Raman data mentioned above.
[0112]
[0113] Among them, S n It is a Raman spectral sequence. This is a resampled spectral sequence.
[0114] Using the above formula (1), a normalized sequence (target Raman data) with all values distributed between -1 and 1 can be obtained. The server then uses the Piecewise Aggregate Approximation (PAA) algorithm to divide the 1024 sequences into multiple new sequences. Each new sequence can be considered as the result of downsampling the target Raman data. In this embodiment, two new sequences are obtained, corresponding to the first downsampled Raman data and the second downsampled Raman data, respectively. Next, the server maps these two downsampled sequences to polar coordinates to obtain the radius and angle values represented by each point. This step can be achieved using the following formula (2).
[0115]
[0116] in, These are spatial angle values in polar coordinates. Let t be the radius in polar coordinates. i Let B be the timestamp of the i-th group, and B be a constant factor for the span of the regularized polar coordinate system.
[0117] After obtaining this data, the Gram angle field can consider the trigonometric sum between each data point from the perspective of cosine angles, thereby identifying the temporal correlation between different groups. Finally, the correlation values between each data point form the final Raman image, a process achieved by the following formula (3).
[0118]
[0119] in, This is a Raman image.
[0120] 303. The server obtains the historical physiological data of the target object, which is used to record the preset state that the target object was in.
[0121] In the case of a patient, this historical physiological data is also referred to as the patient's past medical history or medical history information. Correspondingly, the preset state is also known as the reference disease; the preset state that the target subject was previously in refers to the reference disease the patient previously suffered from. In this embodiment, the reference diseases include five types: percutaneous coronary intervention (PCI), essential hypertension (EH), diabetes mellitus (DM), acute ischemic stroke (ACI), and smoking (SM). Research and analysis have shown that all five reference diseases are associated with identifying subtypes of cardiovascular disease.
[0122] In one possible implementation, the server obtains the target object's historical physiological data from the target object's electronic medical record.
[0123] The acquisition and use of the electronic medical record were done with the consent of the target individual.
[0124] 304. The server determines the preset state that the target object was in based on historical physiological data.
[0125] In the case where the target is a patient, the preset state that the target was in refers to the reference disease that the patient previously suffered from.
[0126] 305. The server determines the statistical probability matrix corresponding to the preset state. The statistical probability matrix is used to reflect the probability that the target object is in multiple candidate states.
[0127] The statistical probability matrix is based on statistics and can reflect the correlation between being in multiple candidate states and other factors to a certain extent. In this embodiment, the statistical probability matrix is related to the historical physiological data of the target object. The historical physiological data is used to record the preset state that the target object was in. Accordingly, the statistical probability matrix can reflect the correlation between the preset state that the target object was in and being in multiple candidate states.
[0128] The method for determining this statistical probability matrix includes the following steps.
[0129] For any preset state among multiple preset states, determine the number of objects that have previously been in that preset state, and then determine the number of objects that are currently in each of the multiple candidate states. For any candidate state among the multiple candidate states, divide the number of objects that have previously been in that preset state and are also in that candidate state by the total number of objects that have previously been in that preset state to obtain the statistical probability of that candidate state in that preset state. Repeat the above steps to obtain the statistical probabilities of other candidate states in that preset state, and then obtain the statistical probabilities of multiple candidate states in multiple preset states, finally obtaining state statistics information. This state statistics information can reflect the correlation between multiple candidate states and multiple preset states. Based on the preset states that the target object has previously been in, the server obtains the statistical probability matrix corresponding to that preset state from the state statistics information.
[0130] Taking patients as the target group as an example, we determine the number of patients with a specific subtype among multiple patients who have had each reference disease. Then, dividing this number by the total number of patients corresponding to that medical history yields the proportion of a single medical history for each subtype. All these calculated probability proportions constitute the state statistics, which represent the degree of influence of each medical history on the diagnosis of each subtype. It is important to note that only medical history information from the training dataset can be used.
[0131] In one possible implementation, the server obtains state statistics, which include the probability of being in multiple candidate states given previous preset states. The server then obtains the statistical probability matrix corresponding to the preset state from the state statistics.
[0132] 306. The terminal performs multi-scale feature extraction on the first Raman image and the second Raman image to obtain multi-scale Raman features.
[0133] Multi-scale feature extraction enables feature extraction from the first and second Raman images across different dimensions, resulting in multi-scale Raman features with stronger expressive power. Feature extraction aims to find the most concise and informative embedding vectors to enhance representation performance. Furthermore, it is used to extract features from the original signal, enabling the algorithm to achieve stronger discriminative capabilities. In some embodiments, step 306 can be implemented by the multi-scale feature extraction module of an object classification model.
[0134] In one possible implementation, the server performs convolution processing on the first Raman image using a first convolution kernel to obtain first Raman image features. The server then performs convolution processing on the second Raman image using a second convolution kernel to obtain second Raman image features, wherein the first and second convolution kernels have different sizes. The server then fuses the first and second Raman image features to obtain the multi-scale Raman features.
[0135] In Raman images, high resolution preserves more local cosine relationships between points, while low resolution expresses more global semantic connections between different groups. With the development of deep learning, convolutional neural networks have become the most popular image analysis method in computer vision. Given that the first Raman image has a resolution of 32 and the second Raman image has a resolution of 64, the size of the first convolutional kernel is 3 and the size of the second convolutional kernel is 5, thus maintaining a rough balance of receptive fields when performing convolutions on the first and second Raman images.
[0136] In this implementation, the server can process the first Raman image and the second Raman image using convolution kernels of different sizes to achieve multi-scale feature extraction and finally obtain multi-scale Raman features, which can more accurately represent Raman data.
[0137] For example, the server slides a first convolution kernel across the first Raman image, performing convolution operations at the locations covered by the kernel to obtain the first Raman image features. The server then slides a second convolution kernel across the second Raman image, performing convolution operations at the locations covered by the kernel to obtain the second Raman image features. The server then concatenates the first and second Raman image features to obtain the multi-scale Raman feature.
[0138] For example, the above process can be achieved by the following formula (4).
[0139]
[0140] in, The first Raman image feature, For the second Raman image feature, G f1 and G f2 For feature extractors in convolutional neural networks, θ f1 and θ f2 These are the first convolution kernel and the second convolution kernel, respectively.
[0141] 307. The terminal performs classification based on the multi-scale Raman features to obtain the first classification result of the target object.
[0142] Among them, classification based on the multi-scale Raman features, that is, classification of target objects based on Raman images, the first classification result can also be called the Raman image classification result.
[0143] In one possible implementation, the server performs a fully connected operation on the multi-scale Raman features to obtain a first classification result for the target object, which includes the possibility that the target object is in multiple candidate states.
[0144] In this implementation, the multi-scale Raman features can be directly mapped to the first classification result through a fully connected layer, resulting in high processing efficiency.
[0145] For example, the server multiplies the multi-scale Raman feature with a fully connected matrix to obtain the first classification result of the target object. For instance, the server performs a fully connected operation on the multi-scale Raman feature using the following formula (5) to obtain the first classification result.
[0146]
[0147] in, This first classification result, also known as the preliminary classification result based on Raman image prediction, is G. l For predictors based on fully connected neural networks, For multi-scale Raman features, It is to use the features of the first Raman image Second Raman image features Obtained by splicing, θ l It is a fully connected matrix.
[0148] The following example uses patients as the target audience and patient classification as the application scenario, combining... Figure 4 The above steps 302, 306 and 307 will be explained.
[0149] See Figure 4 The server downsamples the raw Raman spectrum (Raman data). During downsampling, a first sampling rate is used to obtain first downsampled Raman data, and a second sampling rate is used to obtain second downsampled Raman data. The first and second downsampled Raman data are processed using the Gram angle field algorithm to obtain first and second Raman images. Multi-scale feature extraction is performed on the first and second Raman images using a convolutional neural network to obtain multi-scale Raman features. Classification is then performed based on these multi-scale Raman features to obtain the first classification result for the target object, which is a preliminary subtype classification based on the image data.
[0150] 308. Based on the first classification result and the statistical probability matrix, the terminal determines the second classification result of the target object. The second classification result is the target state of the target object, and the target state belongs to the multiple candidate states.
[0151] The second classification result is obtained based on the first classification result and the statistical probability matrix. The first classification result is a Raman image classification result, and the statistical probability matrix is equivalent to a statistical classification result. The second classification result integrates the two results, resulting in higher accuracy. In some embodiments, step 308 can be implemented by the multimodal data fusion module of the object classification model.
[0152] In one possible implementation, the server concatenates the first classification result with the statistical probability matrix to obtain a multimodal matrix. The server then multiplies this multimodal matrix with the weight matrix to obtain a prediction matrix, which has the same size as the weight matrix. Based on this prediction matrix, the server determines the second classification result for the target object.
[0153] In the field of information representation, the term "modality" refers to a medium that can transmit information. Therefore, "multimodality" can be defined as a method that combines multiple media in information transmission, communication, and representation. In other words, multimodality allows various forms of information representation to interact closely within a system. In this embodiment, medical history and Raman images are combined in a multimodal manner to improve subtype classification performance. Each weight in the weight matrix represents the mixing ratio of the first classification result and the statistical probability matrix. This weight matrix is an adaptive weight matrix, which is randomly initialized in the initial stage and the final weight matrix is obtained through a continuous iterative process. Multiplying the multimodal matrix by the weight matrix yields the prediction matrix, which can be represented by the following formula (6).
[0154]
[0155] in, For the prediction matrix, M is a multimodal matrix. W This is the weight matrix.
[0156] In this implementation, the first classification result and the statistical probability matrix are concatenated to obtain a multimodal matrix. Multiplying the weight matrix by the multimodal matrix yields a prediction matrix, which integrates relevant information from both the image and statistics. Classification based on this prediction matrix results in a highly accurate second classification result.
[0157] To illustrate the above implementation methods more clearly, the method by which the server determines the second classification result of the target object based on the prediction matrix will be described below.
[0158] In one possible implementation, the server sums the values in the prediction matrix according to the dimensions of preset states to obtain a classification matrix. The server normalizes this classification matrix to obtain a state probability distribution of the target object, which includes the probabilities of the target object being in multiple candidate states. Based on this state probability distribution, the server determines the second classification result for the target object. For example, the server determines the candidate state corresponding to the highest probability in the state probability distribution as the second classification result for the target object. The normalization can be implemented using the Softmax function.
[0159] The following example uses patients as the target audience and patient classification as the application scenario, combining... Figure 5 The above step 308 will be explained.
[0160] See Figure 4The first classification result (preliminary subtype classification based on Raman images) and the statistical probability matrix (probability matrix generated based on medical history data) are combined to obtain a multimodal matrix. A weight matrix with adjustable training parameters is used. Multiplying the multimodal matrix by the weight matrix, i.e., allocating the proportions of the Raman prediction result and the probability matrix, yields the prediction matrix. Based on the prediction matrix, a second classification result is determined, which includes the influence degree of each preset state and the subtype of cardiovascular disease (target state).
[0161] It should be noted that the technical solution provided in the embodiments of this application can be implemented by an object classification model, also known as the M3S model. For the process of implementing the technical solution provided in the embodiments of this application using the M3S model, please refer to [link to relevant documentation]. Figure 6 The object classification model uses the content described in steps 301-308 above to classify objects during training, obtains a second classification result, and then performs iterative training based on the difference between the actual classification result and the second classification result.
[0162] In the context of classifying cardiovascular disease patients, the M3S model can combine Raman data and medical history information to complete the classification and diagnosis of cardiovascular disease subtypes. The M3S model combines Raman data and medical history information, driven by multi-scale feature extraction and multi-modal data fusion modules. The probability matrix represents the correlation between medical history and cardiac subtype, while the weight matrix indicates the relative importance of Raman data and medical history information during the decision-making process.
[0163] Accordingly, this application proposes a novel M3S model with two modules, which utilizes Raman spectroscopy and medical history data to realize the diagnostic process of cardiac subtypes. Specifically, the Gram angle field algorithm is used to convert Raman spectra into Raman images of various resolutions. The M3S model includes a multi-scale convolutional neural network structure for feature extraction and representation. A probability matrix is used to represent the interaction between medical history and subtype, and a weight matrix is created to allocate the mixing ratio of multimodal data.
[0164] During the experiments, the technical solutions provided in this application provided reliable experimental evidence, demonstrating that the M3S model exhibits superior performance compared to other state-of-the-art methods in the classification of cardiac subtypes. In summary, it shows that combining two-dimensional Raman images and medical history is a promising method for accurately diagnosing cardiac subtypes. Compared to traditional machine learning and deep learning methods, the addition of multimodal data fusion technology significantly improves the model's subtype classification ability even in one-dimensional spectral data. Table 1 illustrates the technological advancements and specific results of the method provided in this application compared to other methods.
[0165] Table 1
[0166]
[0167] To verify the effectiveness of the model provided in this application, the performance of the M3S model was compared with state-of-the-art algorithms. As shown in Table 1, five randomized tests were performed on all quantitative experiments, and the results were compared by averaging them. The results show that M3S consistently outperforms most of the listed models in its classification performance. This result is reasonable because our model can distinguish subtle features and generate subtype-specific representations, effectively improving its classification performance on the cardiac Raman spectroscopy dataset.
[0168] For traditional machine learning methods, after principal component analysis (PCA), the first 32 principal components are more representative of the original dataset than other dimensions, thus serving as input variables. Based on these components, traditional machine learning methods can classify the major subtypes of heart disease, with the PCA-K nearest neighbor algorithm achieving the best classification accuracy. However, spectral data is often affected by fluorescence and background noise, making it difficult for noise-sensitive machine learning algorithms to accurately distinguish between similar spectra. Therefore, these methods cannot meet the needs of practical applications. In contrast, deep learning methods can learn features that are more robust than principal components, thus achieving higher classification accuracy. Compared to traditional methods, these models show significant improvements, indicating that encoding a single Raman spectrum into a two-dimensional spectral image using a transformation algorithm is more suitable as input for deep learning. Although these deep learning methods reduce the impact of noise sensitivity, they lack consideration for learning from multi-scale and multimodal data.
[0169] This application uses an adaptive weight matrix in its embodiments, which is obtained during the training of the M3S model. To further analyze the M3S model, both manually specified and machine-learned weight matrices are discussed, and the weight matrices are visualized. This helps to qualitatively analyze the advantages and significance of the M3S model.
[0170] First, the M3S model sets the weight ratio for combining Raman spectroscopy and medical history information to 9:1, an empirical value given by medical experts. However, experience is not always optimal, and it cannot be determined whether this is the best fixed ratio for fusion. Furthermore, it is difficult to discover the relationship between medical history, Raman data, and subtypes, posing a challenge to the interpretability of diagnostic results. Therefore, a learnable adaptive multimodal fusion weight matrix was designed to allocate the mixing ratio of multimodal data across different medical histories and subtypes. By comparing multiple indicators, the adaptive weight matrix outperformed the fixed weight matrix, achieving a classification accuracy of 93.3%, demonstrating its effectiveness.
[0171] Secondly, visualize the parameters of the adaptive weight matrix and activate them using the Softmax function. For example... Figure 7 As shown, compared to a fixed weight matrix, an adaptive weight matrix automatically adjusts the mixing ratios of different subtypes and medical histories. Specifically, smoking has a more significant impact on the diagnosis of myocardial infarction than the results predicted by Raman spectroscopy, which aligns with clinical reality. Based on the values of this weight matrix, the model can clearly explain to doctors why a patient was diagnosed with myocardial infarction rather than coronary artery disease or atrial fibrillation. Therefore, the adaptive weight matrix can reflect the deep-seated relationship between medical history and subtype, allowing us to understand previously undiscovered influencing factors and thus make more scientific diagnoses. Compared to expert-perceived values, the mixing ratios learned by the model exhibit better subtype classification performance and interpretability. In conclusion, utilizing an adaptive weight matrix combined with Raman spectroscopy and medical history information is of great significance for the classification of cardiac subtypes.
[0172] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0173] The technical solution provided in this application acquires a first Raman image, a second Raman image, and a statistical probability matrix of a target object. The statistical probability matrix represents the probability that the target object is in multiple candidate states. This matrix is determined based on the target object's historical physiological data, which includes preset states the target object has previously been in. The first and second Raman images have different resolutions. Multi-scale feature extraction is performed on the first and second Raman images to obtain multi-scale Raman features. Classification is performed based on these multi-scale Raman features to obtain a first classification result for the target object. Based on the first classification result and the statistical probability matrix, a second classification result is determined for the target object, which is the target state in which the target object is located. Combining images and the statistical probability matrix for target object classification results in high accuracy.
[0174] This application also provides a target object classification device, see [link to relevant documentation]. Figure 8 The device includes: an acquisition module 801, a feature extraction module 802, a first classification module 803, and a second classification module 804.
[0175] The acquisition module 801 is used to acquire a first Raman image, a second Raman image, and a statistical probability matrix of the target object. The statistical probability matrix is used to reflect the probability that the target object is in multiple candidate states. The statistical probability matrix is determined based on the historical physiological data of the target object. The historical physiological data is used to record the preset states that the target object has been in. The first Raman image and the second Raman image have different resolutions.
[0176] The feature extraction module 802 is used to perform multi-scale feature extraction on the first Raman image and the second Raman image to obtain multi-scale Raman features.
[0177] The first classification module 803 is used to classify based on the multi-scale Raman features to obtain the first classification result of the target object.
[0178] The second classification module 804 is used to determine the second classification result of the target object based on the first classification result and the statistical probability matrix. The second classification result is the target state of the target object, and the target state belongs to the multiple candidate states.
[0179] In one possible implementation, the acquisition module 801 is used to acquire Raman data of the target object, the Raman data including Raman signal values corresponding to multiple time points. The Raman data is then upscaled to obtain the first Raman image and the second Raman image. Historical physiological data of the target object is acquired. Based on the historical physiological data, a preset state in which the target object was previously located is determined. The statistical probability matrix corresponding to the preset state is then determined.
[0180] In one possible implementation, the acquisition module 801 is configured to downsample the Raman data using a first sampling rate and a second sampling rate to obtain first downsampled Raman data corresponding to the first sampling rate and second downsampled Raman data corresponding to the second sampling rate, wherein the first sampling rate and the second sampling rate are different. The first downsampled Raman data and the second downsampled Raman data are then transformed to a polar coordinate system to obtain the radius and angle of each data point in the first downsampled Raman data and the second downsampled Raman data in the polar coordinate system. Based on the radius and angle of each data point in the first downsampled Raman data and the second downsampled Raman data in the polar coordinate system, the correlation between every two data points in the first downsampled Raman data and the second downsampled Raman data is determined. Based on the correlation between every two data points in the first downsampled Raman data and the second downsampled Raman data, the first Raman image and the second Raman image are generated.
[0181] In one possible implementation, the acquisition module 801 is used to acquire state statistics, which include the probability of being in multiple candidate states given that the user has previously been in different preset states. The statistical probability matrix corresponding to the preset state is then obtained from the state statistics.
[0182] In one possible implementation, the feature extraction module 802 is used to convolve the first Raman image using a first convolution kernel to obtain first Raman image features. A second Raman image is then convolved using a second convolution kernel to obtain second Raman image features, wherein the first and second convolution kernels have different sizes. The first and second Raman image features are then fused to obtain the multi-scale Raman features.
[0183] In one possible implementation, the first classification module 803 is used to perform a fully connected operation on the multi-scale Raman features to obtain a first classification result for the target object, which includes the possibility that the target object is in multiple candidate states.
[0184] In one possible implementation, the second classification module 804 is used to concatenate the first classification result and the statistical probability matrix to obtain a multimodal matrix. The multimodal matrix is then multiplied by the weight matrix to obtain a prediction matrix, which has the same size as the weight matrix. Based on the prediction matrix, a second classification result for the target object is determined.
[0185] In one possible implementation, the second classification module 804 is used to sum the values in the prediction matrix according to the dimensions of preset states to obtain a classification matrix. The classification matrix is then normalized to obtain a state probability distribution column for the target object, which includes the probabilities of the target object being in multiple candidate states. Based on this state probability distribution column, a second classification result for the target object is determined.
[0186] In one possible implementation, the second classification module 804 is used to determine the candidate state corresponding to the highest probability in the state probability distribution column as the second classification result of the target object.
[0187] It should be noted that the object classification device provided in the above embodiments is only illustrated by the division of the above functional modules when classifying target objects. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the object classification device and the object classification method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0188] The technical solution provided in this application acquires a first Raman image, a second Raman image, and a statistical probability matrix of a target object. The statistical probability matrix represents the probability that the target object is in multiple candidate states. This matrix is determined based on the target object's historical physiological data, which includes preset states the target object has previously been in. The first and second Raman images have different resolutions. Multi-scale feature extraction is performed on the first and second Raman images to obtain multi-scale Raman features. Classification is performed based on these multi-scale Raman features to obtain a first classification result for the target object. Based on the first classification result and the statistical probability matrix, a second classification result is determined for the target object, which is the target state in which the target object is located. Combining images and the statistical probability matrix for target object classification results in high accuracy.
[0189] The electronic device provided in this application embodiment can be implemented as a terminal or a server. The structure of the terminal will be described first below.
[0190] Figure 9 A structural block diagram of the terminal 900 provided in an embodiment of this application is shown.
[0191] The terminal 900 can be a portable mobile terminal, such as a smartphone, tablet, laptop, or desktop computer. The terminal 900 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other names.
[0192] Typically, terminal 900 includes a processor 901 and a memory 902.
[0193] Processor 901 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 901 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 901 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 901 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 901 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0194] The memory 902 may include one or more computer-readable storage media, which may be non-transitory. The memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 902 are used to store at least one program code, which is executed by the processor 901 to implement the target object classification method provided in the method embodiments of this application.
[0195] In some embodiments, the terminal 900 may also optionally include a peripheral device interface 903 and at least one peripheral device. The processor 901, memory 902, and peripheral device interface 903 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 903 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 904, a display screen 905, a camera assembly 906, an audio circuit 907, a positioning assembly 908, and a power supply 909.
[0196] Peripheral device interface 903 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 901 and memory 902. In some embodiments, processor 901, memory 902 and peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 901, memory 902 and peripheral device interface 903 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0197] The radio frequency (RF) circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 904 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 904 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 904 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0198] Display screen 905 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 905 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 901 for processing. In this case, display screen 905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 905, disposed on the front panel of terminal 900; in other embodiments, there may be at least two display screens 905, disposed on different surfaces of terminal 900 or in a folded design; in other embodiments, display screen 905 may be a flexible display screen, disposed on a curved or folded surface of terminal 900. Furthermore, display screen 905 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 905 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).
[0199] The camera assembly 906 is used to acquire images or videos. Optionally, the camera assembly 906 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 906 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cool light flash, which can be used for light compensation at different color temperatures.
[0200] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 901 for processing, or to the radio frequency circuit 904 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 900. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 901 or the radio frequency circuit 904 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 907 may also include a headphone jack.
[0201] The positioning component 908 is used to determine the current geographic location of the terminal 900 in order to enable navigation or LBS (Location Based Service). The positioning component 908 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, or Russia's Galileo system.
[0202] Power supply 909 is used to supply power to the various components in terminal 900. Power supply 909 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 909 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0203] In some embodiments, the terminal 900 further includes one or more sensors 910. The one or more sensors 910 include, but are not limited to: an accelerometer 911, a gyroscope 912, a pressure sensor 913, a fingerprint sensor 914, an optical sensor 915, and a proximity sensor 916.
[0204] Accelerometer 911 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal 900. For example, accelerometer 911 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 901 can control display screen 905 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 911. Accelerometer 911 can also be used for games or for acquiring user motion data.
[0205] The gyroscope sensor 912 can detect the orientation and rotation angle of the terminal 900. The gyroscope sensor 912, in conjunction with the accelerometer sensor 911, can collect the user's 3D movements on the terminal 900. Based on the data collected by the gyroscope sensor 912, the processor 901 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0206] The pressure sensor 913 can be disposed on the side bezel of the terminal 900 and / or the lower layer of the display screen 905. When the pressure sensor 913 is disposed on the side bezel of the terminal 900, it can detect the user's grip signal on the terminal 900, and the processor 901 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 913. When the pressure sensor 913 is disposed on the lower layer of the display screen 905, the processor 901 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 905. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0207] The fingerprint sensor 914 is used to collect the user's fingerprint. The processor 901 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 914, or vice versa. When the user's identity is identified as trusted, the processor 901 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 914 can be located on the front, back, or side of the terminal 900. When the terminal 900 has physical buttons or a manufacturer's logo, the fingerprint sensor 914 can be integrated with the physical buttons or manufacturer's logo.
[0208] An optical sensor 915 is used to collect ambient light intensity. In one embodiment, the processor 901 can control the display brightness of the display screen 905 based on the ambient light intensity collected by the optical sensor 915. Specifically, when the ambient light intensity is high, the display brightness of the display screen 905 is increased; when the ambient light intensity is low, the display brightness of the display screen 905 is decreased. In another embodiment, the processor 901 can also dynamically adjust the shooting parameters of the camera assembly 906 based on the ambient light intensity collected by the optical sensor 915.
[0209] The proximity sensor 916, also known as a distance sensor, is typically located on the front panel of the terminal 900. The proximity sensor 916 is used to detect the distance between the user and the front of the terminal 900. In one embodiment, when the proximity sensor 916 detects that the distance between the user and the front of the terminal 900 is gradually decreasing, the processor 901 controls the display screen 905 to switch from a screen-on state to a screen-off state; when the proximity sensor 916 detects that the distance between the user and the front of the terminal 900 is gradually increasing, the processor 901 controls the display screen 905 to switch from a screen-off state to a screen-on state.
[0210] Those skilled in the art will understand that Figure 9 The structure shown does not constitute a limitation on terminal 900, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0211] The electronic device provided in this application embodiment can also be implemented as a server. The structure of the server is described below.
[0212] Figure 10 This is a schematic diagram of a server structure provided in an embodiment of this application. The server 1000 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1001 and one or more memories 1002. The one or more memories 1002 store at least one line of program code, which is loaded and executed by the one or more processors 1001 to implement the methods provided in the various method embodiments described above. Of course, the server 1000 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1000 may also include other components for implementing device functions, which will not be elaborated upon here.
[0213] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including program code that can be executed by a processor to perform the classification method for the target objects in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0214] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware, or by a program or program code related to hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0215] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for classifying target objects, characterized in that, The method includes: A first Raman image, a second Raman image, and a statistical probability matrix of a target object are acquired. The statistical probability matrix is used to reflect the probability that the target object is in multiple candidate states. The statistical probability matrix is determined based on the historical physiological data of the target object. The historical physiological data is used to record the preset states that the target object has been in. The first Raman image and the second Raman image have different resolutions. Multi-scale feature extraction is performed on the first Raman image and the second Raman image to obtain multi-scale Raman features; Based on the multi-scale Raman features, a first classification result for the target object is obtained; Based on the first classification result and the statistical probability matrix, a second classification result of the target object is determined. The second classification result is the target state of the target object, and the target state belongs to the multiple candidate states. The classification based on the multi-scale Raman features to obtain the first classification result of the target object includes: A fully connected layer is applied to the multi-scale Raman features to obtain a first classification result of the target object, which includes the possibility that the target object is in multiple candidate states. And, the determination of the second classification result of the target object based on the first classification result and the statistical probability matrix includes: The first classification result and the statistical probability matrix are concatenated to obtain a multimodal matrix; The multimodal matrix is multiplied by the weight matrix to obtain the prediction matrix, wherein the multimodal matrix and the weight matrix have the same size. The prediction matrix is summed according to the dimensions of the preset states to obtain the classification matrix; The classification matrix is normalized to obtain the state probability distribution column of the target object, which includes the probability that the target object is in multiple candidate states; Based on the state probability distribution, the second classification result of the target object is determined.
2. The method according to claim 1, characterized in that, The acquisition of the first Raman image, the second Raman image, and the statistical probability matrix of the target object includes: Acquire Raman data of the target object, the Raman data including Raman signal values corresponding to multiple time points; increase the dimensionality of the Raman data to obtain the first Raman image and the second Raman image; Obtain historical physiological data of the target object; based on the historical physiological data, determine a preset state that the target object was previously in; determine the statistical probability matrix corresponding to the preset state.
3. The method according to claim 2, characterized in that, The step of upscaling the Raman data to obtain the first Raman image and the second Raman image includes: The Raman data is downsampled using a first sampling rate and a second sampling rate to obtain first downsampled Raman data corresponding to the first sampling rate and second downsampled Raman data corresponding to the second sampling rate, wherein the first sampling rate and the second sampling rate are different. The first downsampled Raman data and the second downsampled Raman data are transformed into polar coordinates to obtain the radius and angle of each data point in the first downsampled Raman data and the second downsampled Raman data in polar coordinates. Based on the radius and angle of each data point in the first downsampled Raman data and the second downsampled Raman data in the polar coordinate system, the correlation between every two data points in the first downsampled Raman data and the second downsampled Raman data is determined; The first Raman image and the second Raman image are generated based on the correlation between every two data points in the first downsampled Raman data and the second downsampled Raman data.
4. The method according to claim 3, characterized in that, Determining the statistical probability matrix corresponding to the preset state includes: Obtain state statistics information, which includes the probability of being in the multiple candidate states given that the person has been in different preset states in the past. Obtain the statistical probability matrix corresponding to the preset state from the state statistics information.
5. The method according to claim 1, characterized in that, The multi-scale feature extraction of the first Raman image and the second Raman image to obtain the multi-scale Raman features includes: The first Raman image is convolved using a first convolution kernel to obtain the first Raman image features; The second Raman image is convolved using a second convolution kernel to obtain the features of the second Raman image. The first and second convolution kernels have different sizes. The first Raman image features and the second Raman image features are fused to obtain the multi-scale Raman features.
6. The method according to claim 1, characterized in that, The second classification result of the target object determined based on the state probability distribution column includes: The candidate state with the highest probability in the state probability distribution column is determined as the second classification result of the target object.
7. A classification device for target objects, characterized in that, The device includes: The acquisition module is used to acquire a first Raman image, a second Raman image, and a statistical probability matrix of a target object. The statistical probability matrix is used to reflect the probability that the target object is in multiple candidate states. The statistical probability matrix is determined based on the historical physiological data of the target object. The historical physiological data is used to record the preset states that the target object has been in. The first Raman image and the second Raman image have different resolutions. The feature extraction module is used to perform multi-scale feature extraction on the first Raman image and the second Raman image to obtain multi-scale Raman features; The first classification module is used to classify based on the multi-scale Raman features to obtain the first classification result of the target object; The second classification module is used to determine the second classification result of the target object based on the first classification result and the statistical probability matrix. The second classification result is the target state of the target object, and the target state belongs to the multiple candidate states. Specifically, the first classification module is used for: A fully connected layer is applied to the multi-scale Raman features to obtain a first classification result of the target object, which includes the possibility that the target object is in multiple candidate states. Furthermore, the second classification module is specifically used for: The first classification result and the statistical probability matrix are concatenated to obtain a multimodal matrix; The multimodal matrix is multiplied by the weight matrix to obtain the prediction matrix, wherein the multimodal matrix and the weight matrix have the same size. The prediction matrix is summed according to the dimensions of the preset states to obtain the classification matrix; The classification matrix is normalized to obtain the state probability distribution column of the target object, which includes the probability that the target object is in multiple candidate states; Based on the state probability distribution, the second classification result of the target object is determined.