AI Processor

The AI processor optimizes COVID-19 diagnosis efficiency by using a genetic algorithm to allocate calculation programs across multiple cores, addressing the variability in diagnostic efficiency due to computer performance.

JP7699791B2Active Publication Date: 2025-06-30THE PUBLIC UNIV THE UNIV OF AIZU
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2020194733
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-11-24
Publication Date
2025-06-30
Estimated Expiration
2040-11-24

AI Technical Summary

Technical Problem

The efficiency of COVID-19 diagnosis using machine learning models varies significantly depending on the performance of the computer executing the model, leading to inefficient diagnostic procedures in medical settings.

Method used

An AI processor with multiple arithmetic cores is designed to divide and allocate calculation programs associated with neurons in a CNN model using a genetic algorithm, optimizing communication costs among cores to enhance diagnostic efficiency.

Benefits of technology

This approach enables more efficient diagnosis using machine learning models by optimizing the allocation of calculation programs across multiple arithmetic cores, thereby reducing communication costs and improving diagnostic speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699791000007
    Figure 0007699791000007
  • Figure 0007699791000008
    Figure 0007699791000008
  • Figure 0007699791000009
    Figure 0007699791000009
Patent Text Reader

Abstract

To provide an AI processor that further efficiently performs diagnosis using a diagnostic model.SOLUTION: An AI processor includes a plurality of arithmetic cores. At least one of the plurality of arithmetic cores executes mapping processing of dividing a calculation program associated with each of a plurality of neurons included in a machine learning model of a CNN (Convolutional Neural Network) having a convolutional layer and a fully connected layer, for assignment to each of the plurality of arithmetic cores. Each of the plurality of arithmetic cores executes the calculation program assigned by the mapping processing. In the mapping processing, the calculation program is assigned to the plurality of arithmetic cores by a genetic algorithm so that a communication cost between the plurality of arithmetic cores is equal to or less than a predetermined threshold value.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an AI processor.

Background Art

[0002] Recently, COVID-19 (hereinafter also referred to as coronavirus infectious disease), caused by the SARS-CoV2 virus (hereinafter, novel coronavirus), has been spreading. A standard method for testing whether a person is infected with the novel coronavirus is, for example, reverse transcription polymerase chain reaction (RT-PCR) using a collected sample from a patient, which has a sensitivity of about 60% to 97%. Another method for testing whether a person is infected with the novel coronavirus is, for example, analysis of an X-ray image of a patient's lungs, which has an accuracy of about 80% to 90%.

[0003] Here, when the above-described X-ray image analysis is performed, a doctor needs to manually diagnose each X-ray image of a patient, which causes inefficient diagnostic procedures.

[0004] Therefore, in each hospital, for example, by having a machine learning model (hereinafter also referred to as a diagnostic model) diagnose whether a patient is infected with the novel coronavirus or whether a patient has pneumonia, efficient diagnostic procedures may be performed for many patients.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Non-Patent Documents

[0006]

Non-Patent Document 1

[0007] However, the diagnosis using the above diagnostic model varies greatly in efficiency depending on the performance of the computer that executes the diagnostic model. Therefore, in medical fields such as hospitals, for example, the development of an AI (Artificial Intelligence) processor capable of more efficiently performing the diagnosis using the above diagnostic model is demanded.

[0008] Therefore, an object of the present invention is to provide an AI processor that can perform diagnosis using a diagnostic model more efficiently.

Means for Solving the Problems

[0009] An AI processor according to an aspect of the present invention has a plurality of arithmetic cores, and at least one of the plurality of arithmetic cores divides a calculation program associated with each of a plurality of neurons included in a machine learning model of a CNN (Convolutional Neural Network) having a convolutional layer and a fully connected layer, and executes mapping processing for allocating the divided calculation programs to each of the plurality of arithmetic cores. Each of the plurality of arithmetic cores executes the calculation program allocated by the mapping processing. In the mapping processing, the calculation program is allocated to the plurality of arithmetic cores by a genetic algorithm so that the communication cost among the plurality of arithmetic cores is equal to or less than a predetermined threshold value.

Advantages of the Invention

[0010] According to an aspect of the present invention, it becomes possible to perform diagnosis using a diagnostic model more efficiently.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Embodiments for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Each embodiment is prepared for better understanding of the present invention. However, such embodiments do not limit the technical scope of the present invention. Also, the scope of the present invention encompasses the claims and equivalents thereof.

[0013] [Conventional Information Processing System] First, the conventional information processing system 200 will be described. FIG. 1 is a diagram for explaining the configuration of the conventional information processing system 200. Hereinafter, the case where there are three hospitals (Hospital 11, Hospital 12, and Hospital 1n) will be described, but the number of hospitals may be other than three.

[0014] In the information processing system 200 shown in FIG. 1, at Hospital 11, a doctor makes a diagnosis by analyzing the X-ray image 21 of a patient. Then, the doctor periodically generates patient information 31 indicating the diagnosis result as to whether each patient is infected with the novel coronavirus (e.g., the number of infected patients, the severity of each infected patient, and the number of deaths, etc.) by using, for example, an information processing device (not shown) set in Hospital 11, and transmits the generated patient information 31 to the information processing device 4 of the government.

[0015] Similarly, in the example shown in FIG. 1, at Hospital 12, a doctor generates patient information 32 indicating the diagnosis result of the X-ray image 22 of a patient, and transmits the generated patient information 32 to the information processing device 4 of the government. Also, at Hospital 1n, a doctor generates patient information 3n indicating the diagnosis result of the X-ray image 2n of a patient, and transmits the generated patient information 3n to the information processing device 4 of the government.

[0016] Here, in the information processing system 200, each of Hospital 11, Hospital 12, and Hospital 1n cannot cooperate with other hospitals. In this case, each doctor cannot efficiently diagnose the X-ray image.

[0017] [Information Processing System in the Present Embodiment] Next, the information processing system 100 in the present embodiment will be described. FIG. 2 is a diagram for explaining the configuration of the information processing system 100 in the present embodiment. FIGS. 3 to 5 are diagrams for explaining examples of user interfaces in the present embodiment.

[0018] In the information processing system 100 shown in FIG. 2, at hospital 11, the X-ray image transmitted from the patient's mobile terminal 51 is input to the diagnostic system 61 installed in hospital 11. The detection system 61 is, for example, an information processing device that executes a diagnostic model for determining whether the X-ray image transmitted from the mobile terminal 51 is an image of a patient infected with the novel coronavirus. Then, the diagnostic system 61 transmits the diagnostic result regarding the X-ray image transmitted from the mobile terminal 51 to the mobile terminal 51.

[0019] Similarly, in the example shown in FIG. 2, the diagnostic system 62 installed in hospital 12 diagnoses the X-ray image transmitted from the mobile terminal 52 and transmits the diagnostic result to the mobile terminal 52. Also, the diagnostic system 6n installed in hospital 1n diagnoses the X-ray image transmitted from the mobile terminal 5n and transmits the diagnostic result to the mobile terminal 5n.

[0020] As a result, at hospitals 11, 12, and 1n, it becomes possible to reduce the burden on doctors associated with diagnosing X-ray images. Also, at hospitals 11, 12, and 1n, it becomes possible to prevent incorrect diagnoses of X-ray images. Also, each doctor can, for example, view the user interface U1 shown in FIG. 3 to confirm the diagnostic status of each patient in real time.

[0021] Also, at hospitals 11, 12, and 1n, by performing diagnosis using the diagnostic model, it becomes possible to quickly diagnose X-ray images and quickly notify the mobile terminals 51, 52, and 5n of the diagnostic results. Therefore, each patient can, for example, view the user interface U2 shown in FIG. 4 to quickly confirm the diagnostic result regarding the X-ray image.

[0022] Then, the diagnosis system 61 generates in real time patient information indicating the diagnosis results by the diagnosis model (for example, the number of infected persons, the severity of each infected person, the number of deaths, etc.), and transmits the generated patient information to the government information processing device 4.

[0023] Similarly, in the example shown in FIG. 2, each of the diagnosis system 62 and the diagnosis system 6n generates in real time patient information indicating the diagnosis results by the diagnosis model and transmits it to the government information processing device 4.

[0024] Specifically, each of the diagnosis system 61, the diagnosis system 62, and the diagnosis system 6n transmits patient information to, for example, the cloud server 7. Then, the cloud server 7 generates, for example, real-time information 8 from the diagnosis system 61, the diagnosis system 62, and the diagnosis system 6n, and transmits the generated real-time information 8 to the government information processing device 4.

[0025] As a result, the hospitals 11, 12, and 1n can quickly transmit the latest patient information to the government information processing device 4. Therefore, a government official can quickly confirm the diagnosis results for each patient, for example, by viewing the user interface U3 shown in FIG. 5.

[0026] [Specific Example of Diagnosis Model] Next, a specific example of the diagnosis model MD executed in the diagnosis system 61, the diagnosis system 62, and the diagnosis system 6n will be described. FIG. 6 is a diagram for explaining a specific example of the diagnosis model MD.

[0027] In the diagnosis model MD shown in FIG. 6, when an X-ray image is input from the input layer, it outputs "Normal", which is a category indicating that the X-ray image is not infected with the novel coronavirus (not having pneumonia), or "Suspect", which is a category indicating that there is a suspicion that the X-ray image is infected with the novel coronavirus (there is a suspicion of having pneumonia).

[0028] Specifically, the diagnostic model MD shown in FIG. 6 performs classification of X-ray images, for example, by using a softmax function. Further, the diagnostic model MD shown in FIG. 6 uses, for example, the cross-entropy loss represented by the following formula (1) as the loss function E.

[0029]

Equation

[0030] In Equation (1), w o and b o are parameters in the diagnostic model MD, and each of y and y' represents an actual label (correct label) and a predicted label, respectively. In the diagnostic model MD shown in FIG. 6, by using Equation (1), the loss function E between the actual label and the predicted label is calculated. Then, in the diagnostic model MD, minimization of the loss function E using a stochastic gradient descent algorithm or a backpropagation algorithm is performed, and further, optimization of the parameters is performed.

[0031] [Update of Diagnostic Model] Next, a process for updating the diagnostic model MD (hereinafter also referred to as an update process) will be described. FIG. 7 is a diagram for explaining the update process in the diagnostic model MD. Specifically, FIG. 7(A) is a flowchart diagram for explaining the update process. Further, FIG. 7(B) is a diagram for explaining the transmission and reception of information during the execution of the update process. Hereinafter, the update process performed between the diagnostic system 61 and the cloud server 7 described in FIG. 2 will be explained.

[0032] The diagnostic system 61 preliminarily generates a diagnostic model MD in Hospital 11 by using a dataset including X-ray images transmitted from patients in Hospital 11. Then, the diagnostic system 61 transmits parameters (for example, the gradient ∇gL shown in the following formula (2)) regarding the generated diagnostic model MD to the cloud server 7 (S1).

[0033]

Number

[0034] Subsequently, in response to receiving parameters from each of the diagnostic systems 61, 62, and 6n, for example, the cloud server 7 calculates a global parameter (e.g., the global gradient ∇gG shown in the following equation (3)) from each of the received parameters (S2).

[0035] Furthermore, the cloud server 7 transmits the calculated global parameter to each of the diagnostic systems 61, 62, and 6n (S3).

[0036]

Number

[0037] Thereafter, in response to receiving the global parameter transmitted from the cloud server 7, the diagnostic system 61 updates the parameters of the diagnostic model MD in Hospital 11 as shown in the following equations (4) and (5) (S4).

[0038]

Number

[0039]

Number

[0040] In equations (4) and (5), Wr and br respectively represent the weights and biases in the processing of S1 to S4 performed at the r-th time (the r-th training round). Also, in equations (4) and (5), η represents the learning rate.

[0041] That is, in the information processing system 100, the diagnostic model MD (parameters of the diagnostic model MD) of each hospital is generated by Federated Model Learning (FML).

[0042] As a result, the information processing system 100 can improve the accuracy of the diagnostic model MD of each hospital without transmitting the dataset including the personal information of the patients in each hospital to the outside such as other hospitals. Therefore, the diagnostic model MD of each hospital can accurately diagnose the novel coronavirus while protecting the privacy of each patient.

[0043] Note that the processes from S1 to S4 may be repeatedly performed until the respective determination accuracies of the diagnostic models MD of each hospital exceed the required conditions.

[0044] Also, the diagnostic model MD of each hospital may be generated in a computer other than, for example, the diagnostic systems 61, 62, and 6n (for example, the host computer HC shown in FIG. 8 described later).

[0045] [Mapping of Diagnostic Model] Next, a process of mapping the diagnostic model MD to a conventional AI processor (hereinafter also referred to as an AI chip) PR (hereinafter also referred to as a mapping process) will be described.

[0046] The AI processor PR is, for example, a processor mounted on the diagnostic systems 61, 62, and 6n. When the mapping of the diagnostic model MD to the AI processor PR is performed, in the AI processor PR, for example, after clustering the neurons constituting the diagnostic model MD into a plurality of groups, mapping of each group to each of the plurality of arithmetic cores included in the AI processor PR is performed.

[0047] However, in the conventional mapping process, the communication cost between multiple arithmetic cores may not be considered. Therefore, in the conventional mapping process, the mapping process may not finish within the required time.

[0048] Also, both the clustering of neurons and the mapping to multiple arithmetic cores are generally NP-hard problems and may not be optimally solved in polynomial time.

[0049] Therefore, in the AI processor PR in this embodiment, by using a genetic algorithm, the mapping of the neurons constituting each layer of the diagnostic model MD is performed so as to suppress the communication cost between multiple arithmetic cores.

[0050] [Specific Example of AI Processor in this Embodiment] Next, a specific example of the AI processor PR in this embodiment will be described. FIG. 8 is a diagram for explaining a specific example of the AI processor PR in this embodiment. Hereinafter, it will be described on the assumption that the diagnostic model MD is generated in the host computer HC.

[0051] The AI processor PR shown in FIG. 8 has 15 arithmetic cores. Specifically, the AI processor PR shown in FIG. 8 has 10 arithmetic cores C associated with the convolutional layer and 3 arithmetic cores F associated with the fully connected layer. Also, each arithmetic core is connected to a router R. Hereinafter, the case where the AI processor PR has 15 arithmetic cores will be described, but the AI processor PR may have a different number of arithmetic cores.

[0052] Also, the AI processor PR shown in FIG. 8 has one arithmetic core U associated with a pooling layer or an activation function, and one arithmetic core I / O that transmits weight coefficients to each arithmetic core. Note that the arithmetic core I / O may perform communication with other AI processors PR (inter-chip communication) or communication with a host computer HC that controls each AI processor PR, for example.

[0053] Also, the AI processor PR shown in FIG. 8 has an on-chip memory M (hereinafter also simply referred to as memory M) that stores an input to the AI processor PR. Each arithmetic core such as arithmetic core C loads the input from the memory M and executes processing. Then, each arithmetic core stores the output of each arithmetic core accompanying the execution of the processing in the memory M in order to enable the execution of processing corresponding to the next layer.

[0054] Furthermore, the AI processor PR shown in FIG. 8 has a router R corresponding to each of the arithmetic core and the memory M, and an External DRAM (Dynamic Random Access Memory).

[0055] [Mapping Processing in the Present Embodiment] Next, the mapping processing in the present embodiment will be described. FIG. 9 is a flowchart for explaining the mapping processing in the present embodiment. FIGS. 10 to 21 are diagrams for explaining the mapping processing in the present embodiment.

[0056] The arithmetic core U of the AI processor PR randomly determines N mapping solutions for neurons constituting the diagnostic model MD (neural network) (S11).

[0057] Specifically, when the AI processor PR has four arithmetic cores that can be associated with eight neurons each, and the number of neurons constituting the diagnostic model MD is 30, the number of combinations of mapping solutions is 496, which is the number of combinations of arranging 30 neurons among 32 neurons. Therefore, in this case, the arithmetic core U randomly determines N mapping solutions corresponding to N of these 496 combinations.

[0058] Hereinafter, it will be described assuming that the AI processor PR has four arithmetic cores that can be associated with eight neurons each, and the number of neurons constituting the diagnostic model MD is 30. Further, it will be described assuming that the diagnostic model MD is a neural network having layers L1, L2, and L3 each consisting of 10 neurons.

[0059] [Specific Example of Mapping Result] Next, a specific example of the mapping result in the present embodiment will be described. FIG. 10 is a diagram showing the number of neurons mapped to each arithmetic core. FIG. 11 is a diagram showing the identification information of the neurons mapped to each arithmetic core.

[0060] The examples shown in FIGS. 10 and 11 indicate that the neurons associated with the arithmetic core in the first row and first column are three neurons (neurons 2, 3, and 5) corresponding to layer L1, two neurons (neurons 11 and 12) corresponding to layer L2, and three neurons (neurons 22, 23, and 28) corresponding to layer L3.

[0061] In addition, the examples shown in FIGS. 10 and 11 indicate that the neurons associated with the arithmetic core in the first row and second column are three neurons (neurons 6, 7, and 8) corresponding to layer L1, one neuron (neuron 13) corresponding to layer L2, and three neurons (neurons 21, 29, and 30) corresponding to layer L3. Note that the examples shown in FIGS. 10 and 11 indicate that the number (F) of neurons that can be further associated with the arithmetic core in the first row and second column is 1.

[0062] In addition, the examples shown in FIGS. 10 and 11 indicate that the neurons associated with the arithmetic core in the second row and first column are two neurons (neurons 1 and 4) corresponding to layer L1, three neurons (neurons 15, 17, and 19) corresponding to layer L2, and three neurons (neurons 24, 25, and 26) corresponding to layer L3.

[0063] In addition, the examples shown in FIGS. 10 and 11 indicate that the neurons associated with the arithmetic core in the second row and second column are two neurons (neurons 9 and 10) corresponding to layer L1, four neurons (neurons 14, 16, 18, and 20) corresponding to layer L2, and one neuron (neuron 27) corresponding to layer L3. Note that the examples shown in FIGS. 10 and 11 indicate that the number (F) of neurons that can be further associated with the arithmetic core in the second row and second column is 1.

[0064] Returning to FIG. 9, the arithmetic core U deletes inappropriate mapping solutions from the N mapping solutions determined in the process of S11 (S12).

[0065] Specifically, for example, as shown in FIG. 12, the arithmetic core U deletes a mapping solution in which there is an arithmetic core (the arithmetic core in the first row and first column) associated with nine neurons. In addition, for example, as shown in FIG. 13, the arithmetic core U deletes a mapping solution in which the same neuron (neuron 8) is associated with a plurality of arithmetic cores.

[0066] Then, the arithmetic core U calculates the communication cost corresponding to each of the mapping solutions not deleted in the process of S12 (S13).

[0067] Specifically, the communication cost of each mapping solution is expressed by the following formula (6).

[0068]

Equation

[0069] In Equation (6), d i,j represents the distance between neuron i and neuron j, and c i,j represents the connection status between neuron i and neuron j.

[0070] Specifically, d i,j is, for example, the value obtained by adding 1 to the number of routers R existing between the arithmetic core associated with neuron i and the arithmetic core associated with neuron j. Also, c i,j is 1, for example, when neuron i and neuron j are directly connected, and 0 when neuron i and neuron j are not directly connected.

[0071] More specifically, in the example described in FIG. 11, neuron 1 and neuron 2 are each associated with layer L1. Therefore, in this case, d 1,2 is 1 and c 1,2 is 0. Also, in the example described in FIG. 11, neuron 1 is associated with layer L1 and neuron 14 is associated with layer L2. Therefore, in this case, d 1,14 is 4 and c 1,2 is 1.

[0072] Subsequently, the arithmetic core U identifies M mapping solutions among the mapping solutions not deleted in the process of S12, for which the communication cost calculated in the process of S13 satisfies the condition (S14).

[0073] Specifically, the arithmetic core U identifies M mapping solutions from the mapping solutions not deleted in the process of S12 in descending order of the communication cost calculated in the process of S13.

[0074] Thereafter, the arithmetic core U determines N - M new mapping solutions by crossing over the M mapping solutions identified in the process of S14 (S15).

[0075] Specifically, for example, when the parent 1 and the parent 2 shown in FIGS. 14 and 15 are included in the M mapping solutions identified in the process of S14, as shown in FIG. 16, the arithmetic core U creates new descendants with a ratio of 50 (%) for each of the parent 1 and the parent 2.

[0076] More specifically, in the example shown in FIG. 14, for the arithmetic core in the first row and first column in the case of the parent 1, three neurons corresponding to the layer L1, two neurons corresponding to the layer L2, and three neurons corresponding to the layer L3 are associated. Also, in the example shown in FIG. 15, for the arithmetic core in the first row and first column in the case of the parent 2, one neuron corresponding to the layer L1, four neurons corresponding to the layer L2, and three neurons corresponding to the layer L3 are associated. Therefore, in this case, for the arithmetic core in the first row and first column in the case of the new descendants, as shown in FIG. 16, two (3 * 0.5 + 1 * 0.5) neurons corresponding to the layer L1, three (4 * 0.5 + 2 * 0.5) neurons corresponding to the layer L2, and three (3 * 0.5 + 3 * 0.5) neurons corresponding to the layer L3 are associated.

[0077] Note that, for example, as shown in FIG. 17, when the number of neurons associated with each arithmetic core includes a decimal, the arithmetic core U may perform adjustment so that the number of neurons associated with each arithmetic core becomes an integer, as shown in FIG. 18.

[0078] Returning to FIG. 9, the arithmetic core U causes mutations in N mapping solutions (the sum of the M mapping solutions identified in the process of S14 and the N - M new mapping solutions determined in the process of S15) (S16).

[0079] Specifically, as shown in FIG. 19, for example, the arithmetic core U identifies the arithmetic core at the first row and first column and the arithmetic core at the second row and second column from among the plurality of arithmetic cores shown in FIG. 18, and further identifies layer L1 and layer L2 from among the plurality of layers constituting the diagnostic model MD. Then, the arithmetic core U identifies, for example, 2, which is the minimum value between the number of neurons associated with layer L1 in the identified arithmetic core at the first row and first column and 3, which is the number of neurons associated with layer L2 in the arithmetic core at the second row and second column. After that, the arithmetic core U subtracts 2, which is the value identified as the minimum value, from 2, which is the number of neurons associated with layer L1 in the arithmetic core at the first row and first column, and further adds 2, which is the value identified as the minimum value, to 3, which is the number of neurons associated with layer L2 in the arithmetic core at the first row and first column.

[0080] Similarly, the arithmetic core U subtracts 2 from 3, which is the number of neurons associated with layer L2 in the arithmetic core at the second row and second column, and further adds 2 to 4, which is the number of neurons associated with layer L1 in the arithmetic core at the second row and second column.

[0081] Then, similar to the process of S12, the arithmetic core U determines whether the N mapping solutions after the process of S16 satisfy the constraints. As a result, if there is a mapping solution that does not satisfy the constraints, the arithmetic core U deletes the existing mapping solution. Further, similar to the process of S13, the arithmetic core U calculates the communication cost of the mapping solutions that were not deleted (S17).

[0082] Thereafter, it is determined whether the optimal cost (the minimum cost) among the communication costs calculated in the process of S17 satisfies a predetermined condition (S18).

[0083] As a result, when it is determined that the optimal cost among the communication costs calculated in the process of S17 satisfies the predetermined conditions, the arithmetic core U ends the mapping process (normal end).

[0084] On the other hand, when it is determined that the optimal cost among the communication costs calculated in the process of S17 does not satisfy the predetermined conditions, the arithmetic core U determines, for example, whether the number of executions of the processes after S12 (i.e., the number of generations) has reached a predetermined number (S19).

[0085] As a result, when it is determined that the number of executions of the processes after S12 has not reached the predetermined number, the arithmetic core U performs the processes after S12 again, for example.

[0086] On the other hand, when it is determined that the number of executions of the processes after S12 has reached the predetermined number, the arithmetic core U ends the neuron mapping (abnormal end), for example.

[0087] As described above, the AI processor PR in the present embodiment has a plurality of arithmetic cores, and at least one of the plurality of arithmetic cores executes a mapping process of dividing the calculation programs associated with each of the plurality of neurons included in the diagnosis model MD having a convolutional layer and a fully connected layer and allocating them to each of the plurality of arithmetic cores.

[0088] And in the AI processor in the present embodiment, each of the plurality of arithmetic cores executes the calculation program allocated by the mapping process.

[0089] Specifically, in the mapping process in the present embodiment, the calculation programs are allocated to the plurality of arithmetic cores by a genetic algorithm so that the communication cost among the plurality of cores becomes equal to or less than a predetermined threshold value.

[0090] More specifically, when the mapping process is successfully completed, the arithmetic core U generates a mapping table (not shown) indicating the result of the mapping process. Subsequently, the arithmetic core U downloads the parameters of the diagnostic model MD from the host computer HC. Further, the arithmetic core U refers to the mapping table and transmits the parameters of the diagnostic model MD to each of the arithmetic cores C and F. Also, the arithmetic core U transmits the mapping table to each router R.

[0091] After that, when input data (e.g., an X-ray image of a patient) is input, for example, the arithmetic core U refers to the mapping table and transmits the input data to each of the arithmetic cores corresponding to the neurons included in the first layer. Then, each router R transmits the output data from the first layer to the arithmetic cores corresponding to the neurons included in the next layer in response to the completion of the processing corresponding to the first layer. Further, each router R repeatedly transmits to the arithmetic cores corresponding to the neurons included in the next layer of the layer to be processed until the processing corresponding to the last layer is completed. Then, each router R stores the output data from the last layer (the output data of the diagnostic model MD) in the DRAM.

[0092] That is, the AI processor PR in the present embodiment performs mapping processing by using a genetic algorithm, which is a simpler algorithm than the conventional method and can obtain results within a predictable time.

[0093] Thereby, the AI processor PR in the present embodiment can suppress the communication cost when mapping each neuron to the arithmetic core. Therefore, the AI processor PR in the present embodiment can perform diagnosis using the diagnostic model more efficiently.

Explanation of Signs

[0094] 4: Information processing device 7: Cloud server 8: Real-time information 11: Hospital 12: Hospital 1n: Hospital 51: Mobile terminal 52: Mobile terminal 5n: Mobile terminal 61: Diagnostic system 62: Diagnostic system 6n: Diagnostic system 100: Information processing system MD: Diagnostic model PR: AI processor

Claims

1. having a plurality of arithmetic cores, at least one of the plurality of arithmetic cores executes mapping processing for dividing a calculation program associated with each of a plurality of neurons included in a machine learning model of a CNN (Convolutional Neural Network) having a convolutional layer and a fully connected layer, and allocating the divided calculation programs to respective ones of the plurality of arithmetic cores, each of the plurality of arithmetic cores executes the calculation program allocated by the mapping processing, in the mapping processing, the calculation program is allocated to the plurality of arithmetic cores by a genetic algorithm so that communication cost among the plurality of arithmetic cores becomes equal to or less than a predetermined threshold value, and further, parameters in the machine learning model are updated by using global parameters generated from parameters in the machine learning model and parameters in another machine learning model, the machine learning model is a machine learning model that outputs information as to whether a person corresponding to personal information is in a predetermined state in accordance with input of image data including the personal information, in the mapping processing, a first arithmetic core included in the plurality of arithmetic cores generates a mapping table indicating a result of the mapping processing, in the processing of updating parameters in the machine learning model, the first arithmetic core refers to the mapping table and transmits parameters in the machine learning model to a plurality of other arithmetic cores among the plurality of arithmetic cores, thereby updating parameters in the machine learning model, An AI processor characterized by the above.

2. In Claim 1, the image data is image data in which a patient is reflected, the machine learning model is a machine learning model that outputs information as to whether the patient is in the predetermined state in accordance with input of the image data, An AI processor characterized by the above.

3. In Claim 1, the CNN further has a pooling layer, An AI processor characterized by the above.

4. In Claim 1, in the mapping processing, N mapping solutions for the plurality of arithmetic cores with respect to the calculation program are randomly generated, communication cost in a case where each of the N mapping solutions is adopted is calculated, From the N mapping solutions, identify M mapping solutions in ascending order of the calculated communication cost. Generate N - M new mapping solutions by intersecting the M mapping solutions. Cause mutations to occur in the N new mapping solutions including the M mapping solutions and the N - M new mapping solutions. Recalculate the communication cost when each of the N new mapping solutions with the mutation is adopted. Identify a specific mapping solution among the N new mapping solutions for which the recalculated communication cost is the minimum. Determine whether the communication cost of the specific mapping solution is less than or equal to the predetermined threshold. When it is determined that the communication cost of the specific mapping solution is less than or equal to the predetermined threshold, allocate the calculation program to the plurality of arithmetic cores according to the specific mapping solution. An AI processor characterized by the above.

5. In claim 4, In the mapping process, When it is determined that the communication cost of the specific mapping solution is not less than or equal to the predetermined threshold, for the N new mapping solutions, perform the processes of calculating the communication cost, identifying the M mapping solutions, generating the N - M new mapping solutions, causing the mutation to occur, recalculating the communication cost, identifying the specific mapping solution, and determining whether the communication cost is equal to the predetermined threshold again. An AI processor characterized by the above.

Citation Information

Patent Citations

  • Medical image detection method, device and equipment and storage medium

    CN111652863A

  • Computer arrangement method

    JP2003058520A

  • Method for program segmentation and program for implementing it

    JP2004185271A

  • Neural network device

    JP2020046821A

  • Neural network computation device

    JP2020160564A