Control of a chromatographic separation process

WO2026166743A1PCT designated stage Publication Date: 2026-08-13APPLEXION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-01-19
Publication Date
2026-08-13

Smart Images

  • Figure EP2026051118_13082026_PF_FP_ABST
    Figure EP2026051118_13082026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for machine learning a function configured to control a process of chromatographic separation of a mixture comprising at least two compounds in a multi-column system of simulated-moving-bed type. The function comprises at least one model configured to receive as input operation data and to deliver as output a yield lost with respect to a primary yield for one of the compounds of the mixture. The learning method comprises: - obtaining a dataset; and - machine learning the function based on the obtained dataset. The invention also relates to a method of use of such a learned function, to a computer program for executing such methods, to a storage medium for such a program and to a system containing such a medium.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Description

[0002] Title: Control of Chromatographic Separation Process

[0003] 5

[0004] Scope of the invention

[0005] The present invention relates to the field of computer programs and systems, and more specifically to processes, a system and a program for the control of a chromatographic separation process of a mixture.

[0006] Technical background

[0007] Chromatography is a separation technique based on the difference in distribution of compounds in a mixture between a mobile phase and a stationary phase. The compounds are separated by percolating a mobile phase, consisting of a liquid, gaseous, or supercritical solvent, through a device (called a column or cell) filled with a stationary phase. Separation occurs when all or some of the compounds have different percolation rates. This method is used as an analytical technique to identify and quantify the compounds in a mixture. It can also be used as a purification technique.

[0008] Today, systems exist that can perform a continuous chromatographic separation process, allowing for more efficient use of stationary phases and better separation of complex mixtures compared to traditional single-column processes. Examples of such systems include simulated moving bed (SMB) multi-column systems, as described in US patent 8,282,831. This technology is primarily used for the separation and purification of compounds in the chemical, food, and pharmaceutical industries. In this technology, stationary phase movement is simulated by a cyclic sequence of switching between the inlet and outlet points.Examples of such systems include iSMB systems (acronym for "Improved Simulated Moving Bed"), as described in documents EP 0342629 and.

[0009] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINALUS 5,064,539, SSMB systems (acronym for "Sequential Simulated Moving Bed"), as described in document WO 2015 / 104464, Powerfeed systems as described in document US 5,102,553, ModiCon systems as described in document US 7,479,228, VariCol systems as described in documents US 6,136,198, US 6,375,839, US 6,413,419 and US 6,712,973.

[0010] Such systems typically comprise columns connected in series and filled with stationary phase, along with mixing injection points, eluent injection points, extract collection points, and raffinate collection points that cyclically switch between these columns. These injection and collection points define four zones 1, 2, 3, and 4, and the volumes of mobile phase applied to these four zones over a period of time can be adjusted to optimize system tuning.

[0011] 15 One limitation of such systems is that, when not properly adjusted, cross-contamination can occur at the extract and raffinate collection points. This is because adjusting the mobile phase volumes in the four zones affects how the mixture's components are distributed between the extract and the raffinate. To avoid such contamination, 20 one solution is to maximize the mobile phase volume in zone 1 and minimize it in zone 4. However, this maximizes the volume of water used, resulting in significant dilution. Therefore, this solution is not optimal.

[0012] There is therefore a need to provide an improved control solution 25 of a chromatographic separation process of a mixture in a simulated moving bed type multi-column system.

[0013] Summary of the invention

[0014] The invention relates primarily to a computer-implemented method for the automatic learning of a function configured to control a chromatographic separation process of a mixture comprising at least two compounds in a simulated moving-bed multicolumn system. This method for learning the function is hereafter referred to as the "learning method." The system comprises 35 series-connected columns filled with stationary phase and mixing injection points, eluent injection points, extract collection points, and raffinate collection points that cyclically switch between the columns. The system includes zones 1, 2, 3, and 4. Zone 1 is located between the injection point

[0015] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL of eluent and the extract collection point. Zone 2 is located between the extract collection point and the mixture injection point. Zone 3 is located between the mixture injection point and the raffinate collection point. Zone 4 is located between the raffinate collection point and the eluent injection point.

[0016] 5. The function includes at least one artificial intelligence model configured to take operating data as input and to provide as output a lost yield relative to a primary yield for one of the components of the mixture. The learning process includes:

[0017] - obtaining a dataset comprising, for a set of training samples, operating data and associated lost efficiencies; and

[0018] - machine learning of the function from the obtained dataset.

[0019] In some embodiments, the primary yield is obtained by maximizing the volume of zone 1 or minimizing the volume of zone 4.

[0020] In embodiments, the operating data taken as input by at least one model includes parameters including volumes of raffinate and extract.

[0021] In some embodiments, the function includes a first 20 model. The operating data taken as input by the first model includes concentration measurements at different nodes of the system, and / or concentration measurements in the extract and / or the raffinate.

[0022] In some embodiments, concentration measures at different nodes of the system include:

[0023] 25 - at least one concentration value taken in zone 4 and in zone 1; and

[0024] - at least one concentration value taken in zone 2 and zone 3.

[0025] In some embodiments, the training samples include training samples for the following chromatographic separations:

[0026] - several isotherms, for example of the Langmuir, linear and anti-Langmuir type, including but not limited to:

[0027] o a purification of a lactic acid, for example of type 35 Langmuir;

[0028] o a purification of glucose-fructose sugars, for example of the anti-Langmuir type;

[0029] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINALo a purification of glycerol and salts, for example of linear type.

[0030] The training samples may optionally include training samples for different flow velocities in the 5 system.

[0031] In some embodiments, the function includes a second model. The operating data taken as input by the second model include yield measurements of all compounds from the chromatographic separation in the extract and / or the raffinate.

[0032] In some embodiments, the dataset is obtained by simulating the chromatographic separation process.

[0033] In some embodiments, each model has an architecture comprising:

[0034] - at least two hidden layers;

[0035] 15 - a rectified linear unit function (ReLU); and

[0036] - an output layer including a linear function.

[0037] The invention also relates to a method of using a function learned automatically according to the learning method described above. This method of using the function is hereafter referred to as the "method of use." The method of use comprises:

[0038] - obtaining operational data from a chromatographic separation process in a simulated moving bed multicolumn system; and

[0039] - an application of the function to the operating data 25 obtained.

[0040] In some embodiments, the method of use further includes, after the application of the function, a use of the lost yield provided at the output by the function to control the chromatographic separation process.

[0041] The invention also relates to a first computer program and a second computer program comprising instructions which, when these programs are executed on a computer, cause the computer to implement, respectively, the learning process and the usage process as described above. The invention also relates to a third computer program comprising the first and second computer programs.

[0042] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINALL' invention also relates to a computer-readable storage medium on which one of the computer programs as described above is recorded.

[0043] The invention also relates to a system comprising a processor 5 coupled to a memory. The memory has stored one of the computer programs as described above. The system may also include a simulated multi-column moving bed system as described above or be solely a control system, connected to and separate from the simulated multi-column moving bed system.

[0044] The present invention addresses a need expressed in the prior art. More specifically, it provides a solution for improving the control of a chromatographic separation process of a mixture in a simulated moving bed multicolumn system.

[0045] In particular, in the prior art, there is no solution 15 that optimally reduces cross-contamination within the system. In the invention, the learned function automatically provides a rapid and accurate estimate of the yield lost in the system, thus enabling the chromatographic separation process to be fine-tuned.

[0046] 20

[0047] Brief description of the figs

[0048] [Fig. 1] Figure 1 illustrates examples of flowcharts of the learning process and the usage process.

[0049] [Fig. 2], [Fig. 3], Figures 2 and 3 schematically represent examples of SMB systems.

[0050] [Fig. 4], [Fig. 5], Figures 4 and 5 illustrate an example of a model according to the invention.

[0051] [Fig. 6], [Fig. 7], [Fig. 8], [Fig. 9], Figures 6 to 9 illustrate an example of the data set acquisition step of the learning process.

[0052] [Fig. 10] Figure 10 shows an example of a concentration profile.

[0053] [Fig.11.1], [Fig.11.2], [Fig.11.3], Figures 11.1, 11.2 and 11.3 schematically represent the lost efficiency compared to the primary efficiency provided by the function on different examples.

[0054] [Fig. 12], Figure 12 shows an example of a system.

[0055] 35

[0056] Detailed description

[0057] The invention is now described in more detail and in a non-limiting manner in the following description.

[0058] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINALE With reference to the flowchart in Figure 1, the invention relates to a computer-implemented method for the automatic learning of a function configured to control a chromatographic separation process of a mixture comprising at least two compounds in a simulated moving-bed multicolumn system. This method for learning the function is hereafter referred to as the "learning method." The system comprises series-connected columns filled with stationary phase and mixture injection points, eluent injection points, extract collection points, and raffinate collection points that cyclically switch between the columns. The system includes zones 1, 2, 3, and 4. Zone 1 is located between the eluent injection point and the extract collection point. Zone 2 is located between the extract collection point and the mixture injection point.Zone 3 is located between the mixture injection point and the raffinate collection point. Zone 4 is located between the raffinate collection point and the eluent injection point. The function includes at least one model configured to take operating data as input and to provide as output a lost yield relative to a primary yield for one of the mixture components. The learning process includes obtaining a dataset (S11) comprising, for a set of training samples, operating data and associated lost yields. The learning process includes machine learning (S12) of the function from the obtained dataset.

[0059] The learning process improves the control of the 25 chromatographic separation process of the mixture in the simulated moving bed type multi-column system.

[0060] Indeed, the loss yield provided by the learned function allows for the assessment of the level of cross-contamination within the system, and thus the adjustment of the chromatographic separation process settings accordingly. For example, when the loss yield is significant, the volume of mobile phase applied to the different zones of the system can be modified to improve the yield. Conversely, when the loss yield is low, the volume of mobile phase applied to the different zones of the system can remain the same or be reduced, thereby avoiding an unnecessary increase in the volume of eluent used. The learned function thus enables improved settings for the chromatographic separation process within the user's system.

[0061] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL Furthermore, the learned function automatically provides a fast and accurate estimate of the yield loss in the system based on operating data. Using the learned function is notably faster than performing a process simulation. The function thus allows a user, for example, an engineer or operator in charge of system setup, to obtain system performance without tedious effort or complex calculations, and without requiring in-depth knowledge of the chromatography process: the user provides the operating data, and the function automatically calculates the yield loss. Such a process therefore also improves usability.

[0062] Computer-aided implementation of processes

[0063] The learning process and the usage process are implemented by a computer. This means that the steps (or almost all of the 15 steps) of these processes are executed by at least one computer, or some similar system. Thus, the process steps are carried out by the computer, possibly in a fully automatic or semi-automatic manner. In some examples, the triggering of at least some of the process steps can be achieved through interaction between the user and the computer. The level of interaction required between the user and the computer may depend on the intended level of automation and be balanced with the need to respect the user's wishes. In some examples, this level may be defined by the user and / or predefined.

[0064] For example, the steps S21 for obtaining the operating data and S22 for applying the function to the obtained data can be triggered by the user. In this case, the operating procedure may include the user triggering an estimation of the lost yield for at least one of the components in the mixture. The computer may include a screen and be configured to display a graphical interface to the user on this screen, including a specific button to trigger the steps of the operating procedure. The component in question for which the lost yield is being estimated can, for example, be selected by the user at this time. Triggering and selection can be performed in any way. For example, the user can select the component in question from a list displayed on the interface, or enter the name of the component.The triggering process can then include the user interacting with this button.

[0065] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINALL The S21 data acquisition step can include, after user activation, measuring the system's operating data (e.g., using system sensors) and then recording this data (e.g., to computer memory). The measured operating data can, for example, depend on the selected component. Then, in step S22, the function can be applied to the measured and recorded data. Alternatively, the system can be configured to measure this data regularly, for example, by taking measurements at regular intervals (e.g., every second or every minute). In this case, the S21 data acquisition step can be performed continuously, and user interaction can trigger step S22, which applies the function to the latest measured data.

[0066] In other examples, the S22 application step can also be performed continuously. For example, with each new data measurement, the process might include updating the calculated lost yield by applying the function to the new measured data. In this case, the user interaction might consist solely of selecting the compound for which the lost yield is estimated. A typical example of computer implementation of a process involves running the processes with a system adapted for this purpose. The system might include a processor coupled with memory and a graphical user interface (GUI), the memory containing a computer program with instructions for running the processes. The memory might also store a database.Memory can be any hardware suitable for such storage, possibly comprising several distinct physical parts (for example, one for the program, and possibly one for the database).

[0067] In some examples, the learning process can be performed before the application process. For instance, S12 machine learning might include determining the weights of each model. S12 learning of the function might involve training each model of the function (e.g., in parallel or sequentially). S12 learning of each model can be performed in any number of ways. For example, the weights of each model might be determined to minimize a cost function. This cost function might quantify the difference between lost returns predicted by each model over a portion of the training samples (called the scaling function).

[0068] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL (training dataset) and the actual associated lost yields of these training samples. S12 learning can then include validation of the determined weights by applying each model with the determined weights to the other portion of the training samples (called the validation dataset). Validation of the determined weights can include calculating an accuracy score and comparing this score with a desired accuracy threshold, for example. Once validated, the learning process can include storing the determined weights for each model, for example, in memory.In this case, the S22 application of the function may include retrieving the weights stored in memory, and using each model, with the weights retrieved for that model, on the operating data obtained in step S21.

[0069] 15 In other examples, the learning process can be carried out during the use process. In other words, the use process can include learning the function, i.e., steps S11 and S12. In this case, the learned function can be used in step S22 directly after being learned in step S12.

[0070] 20

[0071] Implementation of the processes

[0072] In some examples, the learning process can be carried out during an initial "offline" phase (or training phase). For instance, the learning process can be executed on a computer server comprising memory on which the dataset is stored and a processor that performs the function learning from the dataset. Steps S11 and S12 can then be executed by this server.

[0073] The operating procedure can be carried out during a second phase, known as the "online" phase (or inference phase). For example, the server can be integrated into a system comprising one or more "client" computers, and the operating procedure can be executed within this system by the server and the one or more "client" computers. For example, each "client" computer can be connected to a respective simulated moving-bed multicolumn system in which a chromatographic separation process is underway. In this case, obtaining the operating data (S21) can include a measurement, for the chromatographic separation process underway in the respective system, of one of the

[0074] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL involves so-called "client" computers, operating data, the transmission of this operating data by the "client" computer to the server, and the server's reception of this data. The function can then, in step S22, be applied by the server to the received operating data to predict the lost yield for the process currently underway in the respective system of the "client" computer. The operating procedure can then include the server sending the predicted lost yield to the "client" computer so that a user of the latter can potentially adjust the settings of the current process accordingly. For example, the "client" computer can be configured to display the lost yield to the user.

[0075] General Presentation of the Chromatographic Separation Process The invention relates to the control of a chromatographic separation process for a mixture comprising at least two compounds in a simulated moving bed multicolumn system. This chromatographic separation process is now discussed in more detail. The chromatographic separation process of the mixture is carried out in the simulated moving bed multicolumn system. This system comprises an array of several chromatography columns containing a stationary phase. The chromatographic separation process comprises successively, cyclically, in a given part of the system:

[0076] - a step of collecting a raffinate, a step of injecting the mixture to be separated, a step of collecting an extract and a step of injecting 25 of mobile phase.

[0077] The various steps described above occur sequentially in this part of the system. This part of the system is preferably located between the output of one column and the input of the next (i.e., each injection or collection is preferably performed between two successive columns). Alternatively, this part of the system may include a column or a portion of a column (i.e., each injection or collection may potentially be performed within a single column).

[0078] At any given time, one or more of the above steps can be implemented simultaneously in one or more parts of the system. For example, all of these steps can be implemented simultaneously in their respective parts of the system.

[0079] By "mixture", or "mixture to be separated", we mean a mixture of species (or compounds, including molecules) containing at least

[0080] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL two species, for example at least one species of interest and at least one impurity. The mixture to be separated may be binary, when it is essentially composed of two species, or complex, when it is composed of more than two species. The mixture to be separated may be diluted in a liquid phase, preferably the mobile phase used in the chromatographic process.

[0081] In some embodiments, the mixture to be separated comprises one or more species selected from:

[0082] - a monosaccharide sugar, for example glucose, fructose, deoxyribose, ribose, arabinose, xylose, lyxose, ribulose, xylulose, allose, altrose, galactose, gulose, idose, mannose, talose, psicose, sorbose or tagatose, and / or a polysaccharide sugar, for example a galacto-oligosaccharide, a fructo-oligosaccharide or a wood hydrolysate, and / or - proteins, and / or

[0083] 15 - amino acids, and / or

[0084] - organic acids such as citric acid, and / or - mineral salts, and / or

[0085] - ionized species, and / or

[0086] - alcohols and / or glycols, and / or

[0087] 20 - organic acids from natural or enzymatic or fermentative environments.

[0088] This list is not exhaustive; the invention in its entirety can be carried out on any chemical species to be separated.

[0089] In some embodiments, the mixture to be separated comprises one or more monosaccharides. Preferably, the extract and the raffinate are enriched in different monosaccharides. Advantageously, the monosaccharide comprises 5 or 6 carbon atoms. Preferably, the monosaccharide is selected from glucose, fructose, deoxyribose, ribose, arabinose, xylose, lyxose, ribulose, xylulose, allose, altrose, galactose, gulose, idose, mannose, talose, psicose, sorbose, tagatose, and mixtures thereof. In some embodiments, the mixture to be separated comprises glucose and fructose.

[0090] In this description, the terms "mixture to be separated", "feed", "mixture to be treated", "product to be purified" and "initial mixture" 35 refer to the same thing.

[0091] By "eluent," we mean the mobile phase injected into the chromatography system to move the components along the columns; thus, the terms "mobile phase" and "eluent" refer to the same thing in the

[0092] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL presents this invention. In processes using resins as stationary phases, the mobile phase or eluent is very often composed of water, sometimes with additives of minerals, acids, or bases. By extension, the term "water" also refers to the mobile phase or eluent.

[0093] 5. By "refined," we mean a fraction enriched in species less retained by the stationary phase. In the case of an initial binary mixture, this is the fraction enriched in the least retained species.

[0094] The term "extract" refers to a fraction enriched in species more retained by the stationary phase. In the case of an initial binary mixture, this is the fraction enriched in the most retained species.

[0095] Chromatographic columns are preferably arranged in series and in a closed loop, with an output of one column connected to an input of the next column, and the output of the last column connected to the input of the first column.

[0096] 15 Columns can also be called chromatography “cells”. They can be used in a carousel system, arranged side by side, or arranged one above the other in one or two towers to limit the floor space required.

[0097] 20. The columns may contain a stationary phase, liquid or solid, preferably solid, in particulate form. The eluent may be a fluid in the gaseous, liquid, or even supercritical state, preferably in the liquid state. Injection lines for the mixture to be separated and the eluent are preferably provided at the inlet of the different columns, and collection lines for the extract and 25% raffinate are preferably provided at the outlet of the columns. These injection and collection lines provided at the inlet and outlet of the different columns constitute the injection and collection points of the system. Preferably, these injection and collection lines are connected via connecting lines between two successive columns.

[0098] In some advantageous embodiments, the system also includes elements for sequencing the injection and collection lines; that is, the injection and collection lines move, preferably synchronously. In particular, the sequencing of these injection and collection lines can be periodic, over a system operating cycle.

[0099] 35 In this application, an "operating cycle" or "cycle" means the time it takes for the injection and collection lines to be sequenced until they return to their initial position in the system. At the end of one cycle, the system is again in its initial configuration. The duration

[0100] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL The time between two shifts in all the injection and collection lines corresponds to one period. A cycle generally comprises as many periods as there are columns in the separation loop. Thus, the cycle of a process implemented on an 8-column system consists of 8 periods.

[0101] 5 The movement of the collection lines (of extract and refine) and the injection lines (of feed and mobile phase) in the system is also referred to as line switching in this description. This line switching can be carried out by opening or closing valves in a fluidic system comprising, on the one hand, connecting lines between successive columns, and on the other hand, one or more supply lines connected to the inlet of each column and one or more sampling lines connected to the outlet of each column (for example, such supply lines and sampling lines can be connected to each connecting line between successive columns 15).

[0102] Four zones can generally be defined in a simulated moving bed (SMB) system:

[0103] - Zone 1, located between the eluent injection point and the extract collection point,

[0104] 20 - zone 2 located between the extract collection point and the injection point of the mixture to be separated,

[0105] - zone 3 located between the injection point of the mixture to be separated and the collection point of the raffinate, and

[0106] - zone 4 located between the raffinate collection point and the eluent injection point 25.

[0107] By "volume" we mean the integrated fluid flow rate over one cycle period (or, equivalently, the average fluid flow rate over one period multiplied by the duration of the period). The period is generally defined as the time interval between the movement of an injection or withdrawal line. In SMB processes, the period is a time interval that dictates the movement of the inlet and outlet lines. In sequential processes, the sequences can be regulated by the volume, and thus the period is of variable duration. Therefore, the "zone 1 volume" corresponds to the total flow rate circulating in the system (specifically in at least one or more successive columns) between the eluent injection point and the extract collection point, integrated over one period. The "zone 4 volume" corresponds to the total flow rate circulating in the system (specifically in at least one or more successive columns) between the point of

[0108] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL collects raffinate and eluent injection points, integrated over a period. "Raffinate volume" means the total volume of raffinate collected over a period; "Extract volume" means the total volume of extract collected over a period; "Eluent volume" means the total volume of eluent injected over a period; and "Mixture volume" (or feed) means the total volume of the mixture injected over a period.

[0109] The volumes of the zones can optionally be expressed using the quantity BV (acronym for "Bed Volume"). The BV quantity corresponds to the volume passing through each of the zones divided by the volume of a cell.

[0110] Utilities and roles of zones

[0111] Optimizing the tuning of an SMB-type system may involve modifying the mobile phase volumes in the different zones 15 over a period.

[0112] Each zone plays a well-defined role in the system to distribute the compounds harmoniously. To illustrate these zone roles, the following paragraphs use a simple example based on at least two compounds contained in the charge to be separated, where A is the notation for the least retained compound, and B is the notation for the most retained compound.

[0113] Operating principle of zones 2 and 3

[0114] Compounds A and B enter the system via charge injection. For compound A to arrive purified at the raffinate, it must follow the movement of the mobile phase 25, and compound B must be predominantly blocked by zone 3. For compound B to arrive purified at the extract, compound A must be predominantly accelerated by zone 2.

[0115] Over a period, the mobile phase volume in zones 2 and 3 must therefore be high enough so that compound A moves more than one column and not so high so that compound B moves less than one column.

[0116] In summary: if the volumes in zone 2 and zone 3 are harmoniously defined, then compounds A and B will be collected purified respectively in the raffinate and extract fractions.

[0117] 35 Thus, each compound enters through the load line and is first distributed through zones 2 and 3, which leads to the definition of a primary yield for each compound.

[0118] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINALPrinciple of operation of zone 1

[0119] Compound B is the most retained compound, so it tends to remain in the stationary phase. With the periodic movement of the inlet and outlet points, it will eventually reach the raffinate stage after several five periods. Therefore, a sufficiently large volume of mobile phase is required in zone 1 to advance compound B more than one column per period. Thus, compound B, present in zone 2 and which ends up in zone 1 after the valve switching, will be pushed towards the extract.

[0120] In summary: if the volume in zone 1 is high enough to move compound B towards the extract, then the raffinate is not likely to be contaminated by compound B. Beyond a certain volume value in zone 1, specific to each separation, contamination towards the raffinate has disappeared, a plateau has been reached and increasing the volume of zone 1 will no longer have an impact.

[0121] 15. Operating principle of zone 4

[0122] Compound A is the least retained, so it tends to follow the mobile phase. Therefore, in zone 4, a sufficiently small volume of mobile phase is applied over a given period to prevent product A from advancing more than one column. If this is not the case, product A will be pushed from zone 4 to zone 1, and consequently, product A will pass through zone 1 and be collected in the extract.

[0123] In summary: if the volume in zone 4 is low enough not to move compound A towards the exit of zone 4, then the extract is not at risk of being contaminated by compound B. Below a certain value of 25 volume in zone 4, specific to each separation, contamination towards the extract has disappeared, a plateau has been reached and decreasing the volume of zone 4 will no longer have an impact.

[0124] Conclusions

[0125] In a balanced system, the roles of each zone can be summarized according to the following criteria:

[0126] - Zone 1: elute B - this implies a high volume in zone 1, - Zone 2: elute A without eluting B,

[0127] - Zone 3: restrict B and eliminate A,

[0128] 35 - Zone 4: brake A - this implies a low volume in zone 4.

[0129] The volume of eluent to be injected during each period is equal to the difference between the volume used in zone 1 and that in zone 4. The volume of feedstock to be injected during each period is equal to the difference between the

[0130] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL volume used in zone 3 and that in zone 2. A system is therefore adjusted through the volumes of mobile phase which cross each of these zones over a period and more generally over a cycle.

[0131] 5. Origins of cross-pollution

[0132] For a balanced system, if compound A is primarily destined for the raffinate, losses of compound A to the extract can come from two possible sources: either because compound A has passed through zone 2 (zone 2 leakage) or because compound A has passed through zone 4 (zone 4 leakage; compound A passing through zone 4 can then be partially collected in the extract). Therefore, there are two possible explanations for the presence of compound A in the extract: a low zone 2 volume or a high zone 4 volume.

[0133] Conversely, if compound B is to go mostly to the extract, the losses of compound B to the raffinate come from two possibilities: either because compound B has passed through zone 3 (zone 3 leakage) or because compound B has passed through zone 1 (zone 1 leakage, compound B passing through zone 1 can then be collected in part in the raffinate).

[0134] Therefore, there are two possible explanations for the presence of compound B in the raffinate: a high zone 3 volume or a low zone 1 volume. Adjusting the distribution of A and B between the extract and the raffinate thus requires verifying the impact of modifying the mobile phase volumes of the four zones.

[0135] The most direct conclusion is that, in order to avoid taking risks and simplify the settings, the mobile phase volume should be maximized in zone 1 and minimized in zone 4, which has the impact of maximizing the volume of eluent used, inducing significant dilution and a decrease in the economic performance of the system.

[0136] Therefore, there is an industrial interest in measuring the contributions of different zones to product distribution within a system. To properly adjust a system, it is particularly important to accurately measure how products are distributed within the different zones.

[0137] In summary, each product actually collected at the extract and raffinate is the result of a combination of a primary contribution via zones 2 and 3 and a lost contribution via zones 1 and 4. This leads to the definition of the concepts of effective yields (as experimentally measurable), and the concepts of primary yields (distribution of

[0138] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL products via zones 2 and 3) and the notions of lost yields (distribution of products via zones 1 and 4).

[0139] Simulation of the chromatographic separation process

[0140] 5 In the present application, the invention uses numerical simulation of chromatographic processes to generate datasets that can be used to train artificial intelligence models. Chromatographic process simulation is common practice, and several tools are available for performing it:

[0141] - Commercial software exists.

[0142] - Free and open-source programming languages, or those under licensed license, can be used to solve the equations necessary for the simulation. - Numerous books describe the equations that can be used to perform a simulation.

[0143] 15. The equations to be solved to perform a simulation serve to reproduce the most important phenomena, including:

[0144] - The equilibria of all or part of the compounds in the feed to be purified between the liquid phase and the solid phase. These equilibrium equations are often called isotherms.

[0145] 20 - Dispersion phenomena occurring during the movement of products along the column: axial dispersion, diffusion, adsorption / desorption kinetics.

[0146] Different equations exist for each of these phenomena, and they must be solved over time and space until a stable cyclic state is reached. Simulating a multi-column process requires solving these equations for several columns, with inlet concentrations adjusted depending on whether water or the product to be treated, and / or the output from another column of the process, feeds one or another of the columns. The compositions of the extracts and raffinates from the simulated process are obtained by integrating or averaging the concentrations over the collection time of the fractions.

[0147] The parameters of the simulation equations are determined using experimental results obtained on one or more columns. Thus, a low-volume injection onto a column allows measurement of the average retention of the compounds (characteristic of the liquid / solid phase equilibrium of each compound) and the peak width (characteristic of dispersion phenomena). Tests with larger injection volumes allow observation of whether the retention time decreases with increasing quantity.

[0148] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL injected (Langmuir isotherm) or if it increases (anti-Langmuir isotherm) or if it remains constant (linear isotherm). These isotherm concepts are described in "Preparative Chromatography" by Schmidt-Traub (ISBN 3-527-30643-9 & 978-3-527-32897-7), "Chromatographic Processes" by Nicoud (ISBN 978-1-107-08236-6) and in "Fundamentals of Preparative and Nonlinear Chromatography" by Guiochon (ISBN 978-0-12-3270537-2).

[0149] In summary, to perform a MO simulation of a particular application, it is necessary to have, for each of the N simulated compounds, parameters and equations to calculate the liquid-solid equilibria and the dispersion phenomena: kinetics and / or diffusion.

[0150] Changing one or more of these parameters should be considered as leading to another simulation M1, then M2, then M3, etc.

[0151] Once a simulation is in place, it can be used to look at how process parameters (concentration of species in the product to be treated, speed of the mobile phase, volume of the mobile phase) affect performance.

[0152] In some examples, the associated lost yields of a portion (or all) of the training samples may have been obtained by simulating the chromatographic separation process in the system. The simulation can be configured to take as input data on the chromatographic separation process, such as the components of the mixture or system and the system settings, and provide as output a yield for one or more of the components of the mixture considered, as a function of the settings. The dataset can be obtained in any way.For example, obtaining the S11 dataset may include, for each sample for which the lost yield is obtained by simulation, data from the two-step use of the simulation: a first step to calculate the effective yield with the setting considered for the sample, and a second step to obtain the primary yield, for example by maximizing the volume in zone 1 and minimizing the volume in zone 4, and a determination of the lost yield by calculating the difference between the 35 yields obtained for these two settings.

[0153] Control of

[0154]

[0155] édé de sé

[0156]

[0157] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL In this application, "yield" or "effective yield" means the ratio of the mass of a compound in one of the two collected fractions (extract or raffinate) to the total mass of that compound in both collected fractions combined, or also the ratio of the mass of that compound in one of the two fractions to the mass in the injected charge.

[0158] As explained in the preceding paragraphs, yield is the combination of two components: the material flow from zones 2 and 3, and the material flow from zones 1 and 4 resulting from cross-contamination. Therefore, this application will distinguish between yields with and without cross-contamination. By definition, the "primary yield" in this application refers to the plant's yield if there were no cross-contamination. The effective yield, on the other hand, is the yield taking cross-contamination into account. The effective yield is the yield that can be experimentally measured by taking representative samples of the extracts and raffinates and measuring their respective concentrations. A "lost yield" is the difference between a primary yield and an effective yield.This primary yield can be considered the yield obtainable when cross-contamination is zero for the compound in question, starting from the mixture in question, with the eluent in question, and in the chromatographic system in question. The primary yield can therefore be considered the yield obtained for the compound when the volume of zone 1 is sufficiently large and the volume of zone 4 is sufficiently small. In other words, the primary yield corresponds to that which would be obtained if the current settings were modified by maximizing the volume of zone 1 and minimizing the volume of zone 4 (which would, on the other hand, lead to maximum eluent consumption), while keeping the volumes of zone 2 and zone 3 constant.

[0159] The usage process can be included in a control process for the chromatographic separation process executed in the system, which may include, after application of the automatically learned function, control of the chromatographic separation process in the system, based on the loss yield provided as output by the function. When the function comprises only one model, the loss yield provided as output by the function can be the loss yield provided by that single model.

[0160] When the function includes multiple models, the lost efficiency provided at the output by the function may correspond to the lost efficiency provided by one of the models. For example, the different models may take as input

[0161] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL Different operating data, and the model used can then be the one for which the operating data is provided. Alternatively, if data is provided for several models, the lost efficiency output by the function can be an average of the 5 lost efficiencies output by the different models (or one of them, for example, selected by the user).

[0162] The use of the lost yield provided as output by the function (S23) may include an analysis of the provided lost yield and, based on the analysis results, an adjustment of the system's operating parameters. For example, the operating parameters may include an applied volume in each zone, and the adjustment may involve modifying the applied volume in one or more zones of the system. The system adjustment may, in particular, lead to a decrease in the lost yield within the system, i.e., reducing cross-contamination, meaning the presence in the raffinate of a more retained compound (which is desired to be collected in the extract), or the presence in the extract of a less retained compound (which is desired to be collected in the raffinate). This phenomenon can also be described as the "leakage" of a compound into one of the zones, as will be explained in more detail below.For example, the modification might include an increase in the volume of zone 1 or a decrease in the volume of zone 4, depending on the analysis results. These volumes can be modified by changing the eluent injection rate and / or the mixture injection rate and / or the extract collection rate and / or the raffinate collection rate; and / or by changing the duration of the period.

[0163] 25. The calculation of the lost yield according to the invention and the adjustment of the operating parameters can be carried out automatically or semi-automatically. The modification of the operating parameters to be made can, for example, be automatically deduced from the result of the analysis and automatically applied to the system, for example, after being validated by the user. The method can, for example, include calculating the sign of the estimated lost yield to deduce the location of the compound leakage and adjusting the settings accordingly. For example, a positive lost yield means that the compound is leaking from zone 1, and the method can then automatically increase the volume of zone 1. A negative lost yield means that the compound is leaking from zone 4, and the method can increase the volume of zone 4.If the sign of a lost return indicates the area to be corrected, its absolute value assesses its importance and thus contributes to the assessment of the magnitude of the correction.

[0164] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL. The process may thus include an evaluation of the magnitude of the volume increase in zone 1 or zone 4 (depending on the sign of the lost yield) as a function of the absolute value of the lost yield.

[0165] In some examples, the setting may also depend on the magnitude of the estimated lost efficiency. For instance, the process might involve changing parameters only when the estimated lost efficiency is significant. For example, operating parameters might be changed automatically when the lost efficiency exceeds a predetermined threshold (in absolute value, for example). Alternatively, the user might choose whether or not to change the operating parameters based on the estimated lost efficiency. The process might then include a display of the estimated lost efficiency and a trigger for the user (for example, via a specific button) to change the parameters when deemed appropriate.The process can also in this case 15 suggest the modification of the parameters to the user when the estimated lost yield is significant (for example when it exceeds the predetermined threshold), for example with a particular display of lost yield to the user in this case.

[0166] The method of use provides a particularly effective and accurate estimate of the yield lost in the system, which improves the tuning of the chromatographic separation process, and therefore, ultimately, its productivity and quality.

[0167] Data taken as input by at least one model

[0168] 25. At least one model is configured to take operating data as input. For example, each model can take its own specific operating data as input. "Operating data" refers to data related to the operation of the chromatographic separation process taking place in the system in question. During the execution of the operating process, this data relates to the operation of the chromatographic separation process whose lost yield is to be calculated.

[0169] In examples, the operating data taken as input by at least one model might include system settings. A "setting parameter" is defined as a parameter affecting the operation of the chromatography process within the system. For example, a setting parameter can be changed by the user during the execution of the chromatographic separation process.

[0170] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL Examples of parameters include, among others, the volumes of raffinate and extract used, which are known physical volumes. The volumes of raffinate and extract correspond respectively to the volumes of raffinate and extract fractions obtained over a period.

[0171] 5. In examples, the operating data may include, for at least one model, measurements taken during the execution of the separation process. A "measurement" is defined as a value quantifying a property or characteristic of the fluid flowing at a given point in the system during the execution of the chromatographic separation process. For example, the operating data may include concentration measurements at different nodes of the system, and / or concentration measurements in the extract and / or the raffinate. A "system node" or "observation node" is defined as a freely chosen physical point in the chromatography system. In some embodiments, the observation node is located between the outlet of one column and the inlet of the next column in the system. The concentration measurements may be concentration measurements of the compound in the mixture for which the yield loss is predicted.During training, the concentration measurements of each sample can therefore be for the compound for which the corresponding lost yield is associated. During inference, the concentration measurements taken as input by at least one model can be for the compound for which the lost yield is to be predicted. The concentration of a compound can quantify the proportion of that compound in the mixture. In examples, the concentration measurements might include at least one concentration value taken from zone 4 and zone 1, and at least one concentration value taken from zone 2 and zone 3 (these values ​​being concentration values ​​for the component for which the lost yield is calculated).

[0172] In examples, the operating data may include, for at least one model, yield measurements of several compounds in the mixture in one of the collected fractions, for example, the extract. The yields of the compounds can be deduced from measurements of compound concentrations in the extract and raffinate, as well as from the volumes of extract and raffinate.

[0173] 35 During training, the dataset may include data obtained for a set of different chromatographic separation processes carried out in the same system or in several different systems (e.g., with stationary phases)

[0174] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL (different eluents, different mixtures to be separated, different system sizes). For example, each training sample may correspond to a respective chromatographic separation process performed with respective operating parameters. The dataset may include samples corresponding to processes of different types of chromatographic separations, and to different settings of these processes. For example, the dataset may include training samples for chromatographic separations using Langmuir, linear, and anti-Langmuir isotherms.For example, the dataset may include training samples for one or more purifications of lactic acid (e.g., Langmuir type), one or more purifications of glucose-fructose sugars (e.g., anti-Langmuir type), and / or one or more purifications of glycerol and salts (e.g., linear type).

[0175] 15 The dataset may also include training samples obtained at different flow velocities in the system(s) in which they are performed.

[0176] Result provided by at least one model

[0177] 20. At least one model is configured to provide as output a lost yield for one of the compounds in the mixture. In other words, the model calculates a lost yield for a given compound in the mixture; that is, a value for that given component, which might be different for another component. This given compound can be the one for which measurements 25 are provided as input to the at least one model. For example, when the concentration or yield measurements taken as input by the at least one model are measurements for a certain compound, the lost yield provided as output by the at least one model can be for that compound.

[0178] The lost yield, which corresponds to the difference between this primary yield and the actual yield, thus quantifies the potential yield gain that could be achieved by modifying the process settings. It therefore allows the user to assess the relevance of applying a process setting modification to improve the yield for the compound in question. In other words, the function identifies for each compound whether it is normally recovered as an extract or as a raffinate, and the lost yield indicates the potential additional yield gain by adjusting the volumes in zone 1 or 4. For example, when seeking to obtain product N, the function can provide information on

[0179] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL 80% yield, with the possibility of gaining +8% maximum (the lost yield) if we increase the volume of zone 1.

[0180] In some embodiments, if the yield loss exceeds a predetermined threshold, the volume in zone 1 or 4 is adjusted to increase the effective yield and reduce the yield loss. In some embodiments, if the yield loss is less than a predetermined threshold, no volume adjustment in zone 1 or 4 is made, so as to maintain the effective yield substantially constant and not increase eluent consumption.

[0181]

[0182] With reference to figures 2 to 12, examples of implementation of the invention are now discussed.

[0183] 15 Examples of SMB-type systems

[0184] Figure 2 schematically represents a first example of an SMB 100 system. Such a system is primarily used for chromatographic separation processes, enabling the separation and purification of compounds in the chemical, food, and pharmaceutical industries. It is a continuous process that allows for more efficient use of stationary phases and better separation of complex mixtures compared to traditional single-column processes.

[0185] The SMB 100 system consists of several columns 25 connected in series 121, 123, 125, filled with stationary phase. A movement of the stationary phase is simulated by a cyclic sequence of switching the input and output points

[0186] In such a system, between these columns 121, 123, 125, the following liquids are injected or withdrawn:

[0187] - Load 112 (also called the mixture): the mixture to be separated is introduced at the inlet of one of the columns. This mixture contains at least two compounds to be separated, A and B. In our example, compound A is less retained by the stationary phase than compound B.

[0188] 35 - The eluent 116: a liquid mobile phase is injected to move the compounds along the columns. In some examples, the eluent may be pure water or water with a low acid or base content (<2% by mass).

[0189] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL- Extract 110: fraction containing predominantly the most retained compound (compound B).

[0190] - Rafinate 114: fraction containing predominantly the least retained compound (compound A).

[0191] 5. The injection and withdrawal points define areas which, on the diagram, have the following meaning and numbering:

[0192] - “Zone 1” 131: applies to columns between the eluent injection point 116 and the extract collection point 110.

[0193] - “Zone 2” 132: applies to columns between extract collection point 110 and charge injection point 112.

[0194] - “Zone 3” 133: applies to columns between the charge injection point 112 and the raffinate collection point 114.

[0195] - “Zone 4” 134: applies to columns between the raffinate collection point 114 and the eluent injection point 116.

[0196] 15 The mixture injection point 112, eluent injection point 116, extract collection point 110, and raffinate collection point 114 switch cyclically between the columns. Figure 2 shows the system in a first state 210, then after a switchover 220 (each injection or collection point being shifted one column relative to the first state).

[0197] Figure 3 schematically represents an example of an SSMB system. For an SSMB system, a period can be divided into several subsequences 310, 320, 330, as shown in the figure. The first subsequence 310 is a loop subsequence, in which the fluid circulates in a closed loop 25 within the system, without fluid injection or collection. The second subsequence 320 is an open-loop water / raffinate subsequence, in which eluent is injected and raffinate is collected, without mixture injection or extract collection, and in which there is no fluid flow in zone 4.The third subsequence 330 is a water / extract sequence, in which mixture and eluent injections, and extract and raffinate collections, are carried out, with zero fluid flow in zone 2 and in zone 4.

[0198] In such a system, the mobile phase volume per zone is the sum of the volumes in each of the sub-sequences:

[0199] - The volume of zone 1 is the sum of the volumes in the loop subsequence 310, the water / raffinate subsequence 320 and the water / extract subsequence 330;

[0200] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL- Zone 2 volume is the sum of the volumes in loop subsequence 310 and water / raffinate subsequence 320; - Zone 3 volume is the sum of the volumes in loop subsequence 310, water / raffinate subsequence 320 and water / extract subsequence 330;

[0201] - The volume of zone 4 is that of the loop subsequence 310. As long as there is an injection of eluent, an injection of charge, a collection of extract and raffinate, it is possible to calculate the mobile phase volumes over a period.

[0202] Neural networks and learning

[0203] The function includes at least one model, which can be a neural network. A neural network is a computational model inspired by the functioning of the 15 biological neurons in the human brain. It consists of nodes called neurons, organized into layers (input, hidden, and output), where each neuron in a layer is connected to those in the next layer by weights. Figure 4 illustrates the principle of a 400-digit neuron: the output is a function of the input data and a neuron-specific weighting for each input.

[0204] 20 Figure 5 illustrates the principle of a 500-type neural network model according to the invention. The key components of such a 500-type model are as follows:

[0205] - 501 neurons (or nodes): basic units that perform calculations. Each neuron receives input signals, weights them, sums them, and applies an activation function to produce an output.

[0206] - Layers: neurons are organized into layers:

[0207] - Input layer 510: receives the initial data.

[0208] - Hidden layers 520, 530, 540: perform the transformation and extraction of features.

[0209] - Output layer 550: produces the final result of the network. - Weights: adjustable parameters that determine the strength of the connections between neurons.

[0210] - Activation function: a function applied to the weighted output of a 35-neuron to introduce non-linearity (such as:

[0211] the Sigmoid, ReLU or Tanh functions).

[0212] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL- Training: process of adjusting weights via an optimization algorithm (such as backpropagation) to minimize the error between network predictions and expected values.

[0213] Training a neural network involves adjusting the weights of the connections between neurons to minimize the error between the network's predictions and the actual values. This is achieved by using a training dataset (also called a "training dataset"). Once the weights are determined, the neural network is used to calculate the result based on the input data values.

[0214] Neural networks can be used for various tasks such as image recognition, machine translation, and many other machine learning applications. The effectiveness of learning depends on the quality of the training dataset. Other alternatives to neural networks can be used in the invention, such as:

[0215] - Random Forests (translation from the English "Random Forests"):

[0216] a set of methods based on decision trees. Random forests create multiple decision trees from random subsets of data and features, then aggregate their results to improve accuracy and reduce overfitting.

[0217] Support Vector Machines (SVMs): classification algorithms that seek to find the best hyperplane separating data classes in a high-dimensional space. SVMs are a set of supervised learning techniques designed to solve discrimination and regression problems.

[0218] - k-nearest neighbors (k-NN): a classification or regression algorithm based on data similarity. For a new data point, k-NN finds the k nearest training data points and uses their labels to predict the class or value of the new data point.

[0219] - Linear regression: a basic technique for regression. Linear regression is used for continuous value prediction problems, where the unknown data value uses other known data values.

[0220] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL System Setup

[0221] It is common practice to take samples of the collected fractions either during a cycle or over a period of time. These samples are then analyzed offline to obtain the concentrations of each of the 5 collected samples.

[0222] It is also common practice to use online sensors to measure the compositions of extract and raffinate fractions, as well as the concentrations of compounds circulating in the system. Reference is made in this regard to document WO 2019 / 097182.

[0223] Thus, to adjust the settings of a system, it is possible to use the compositions of the collected fractions to assess the quality of the compound distribution and make adjustments. Since there are two possible origins (zone 2 / 3 or zone 1 / 4), one approach is to proceed by trial and error, that is, by testing different solutions until one that works is found. However, this potentially means having to repeat the adjustment if the first one is unsuccessful, resulting in a loss of yield, efficiency, and time.

[0224] On the other hand, if concentration measurement information between columns is available, then it is possible to evaluate a concentration profile and to determine the existence of a zone 1 or zone 4 leak.

[0225] However, if the concentration profile shows a simultaneous leak from zone 2 and zone 4, it is very difficult to know what the most important action to take is.

[0226] The present invention solves these problems by allowing the source of the leak to be identified directly, and therefore the adjustment solution to be implemented for the system.

[0227] Operating principle

[0228] In one example, the function according to the invention comprises two artificial intelligence models.

[0229] The first model takes the following input parameters:

[0230] - optionally the volumes of raffinate and extract, and - concentration values ​​of a compound (which is the one for which the lost yield is predicted) on different samples, for example:

[0231] at least one concentration value taken in zone 4 and zone 1, preferably a plurality of such values, for example

[0232] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL three such values ​​for each of zones 4 and 1, each value in a zone being obtained at a different node of the system, o at least one concentration value taken from zone 2 and one concentration value taken from zone 3, and 5 o the concentration values ​​at the raffinate and extract.

[0233] The second model takes the following input parameters:

[0234] - the volumes of refiner and extract, and

[0235] - the yields of a plurality of compounds (for example of all compounds) present in the feed to be treated on one of the two collected fractions.

[0236] As an output parameter, each artificial intelligence model calculates the proportion of compound that passes through zone 1 or zone 4. The function therefore allows the same value to be calculated using two different methods.

[0237] The training of these two artificial intelligence models can be based on a dataset using any simulation tool. Naturally, all available experimental data can be added to this dataset. However, experimental data takes one to two days to obtain, while a simulation requires only a few seconds to a few minutes of computation. Using a simulation to obtain training data therefore saves time. A primary advantage of simulation is thus the time saved. A second advantage lies in the fact that all possible situations can be calculated: balanced operation according to the previously stated criteria, as well as slightly or completely unbalanced operation.

[0238] 25 Regardless of the quality of the equations used, the simulation provides good information on behavioral dynamics.

[0239] Simulation software and process parameters

[0240] Figures 6 to 9 illustrate an example of the data acquisition step of the learning process.

[0241] Simulation software can be used during step S11 to obtain the training dataset or consolidate a dataset containing experimental data. Such simulation software is generally composed of several parts. The first part of the software includes a description / parameterization of the separation. This part allows the user to define the equilibrium parameters for each compound between the mobile and stationary phases, as well as parameters of

[0242] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL dispersion or diffusion (examples of such parameters are shown schematically under the reference "MO Simulation" in the figures).

[0243] In a second part, for a given simulation (MO), the user defines the operating parameters to be applied to the process, 5 in particular the volumes of mobile phase in each zone, the composition of the charge to be simulated and the average speed applied to a zone (examples of such simulation parameters are shown schematically under the reference "M0_Pi" in the figures).

[0244] In the third part, for a given "M0_Pi" parameter and a set of volumes in the four zones BV1, BV2, BV3, and BV4, the simulation software calculates, for each compound, the concentrations in the columns at each time step and the average concentrations in the extract and the raffinate. From this data, the simulation software also calculates the distribution yields of all compounds between the extract and the raffinate.

[0245] 15 The simulation software therefore calculates, for a set of parameters i:

[0246] BV1 i, BV2i, BV3i and BV4, the concentration values ​​of each compound in the different zones, the concentration values ​​in the extract and raffinate fractions and the associated yields.

[0247] In other words, the simulation allows us to calculate, for the set 20 of data < BV1, BV2, BV3, BV4 > 610, the resulting dataset < Concentrations of profiles, extract, raffinate >, and from this set, the resulting dataset < Extracted and / or raffinate yields >.

[0248] By convention, the simulation and at least one artificial intelligence model of the function can provide an extract yield as output. The 25% raffinate yield can then be deduced using the following formula:

[0249] Rafinate yield = 100% - extract yield. Alternatively, it is possible to consider the raffinate yield, and in this case conversely deduce the extract yield from that calculated using raffinate (for example, by using a similar formula).

[0250] Determination of actual, primary and lost yields: impact of zones 1 and 4

[0251] Obtaining the training dataset S11 also includes calculating the fraction of each compound passing through zones 1 or 4. To calculate the yield variations from zones 1 and 4, obtaining the dataset S11 may also include calculating the yield lost for each training data point, for example with any of the following solutions.

[0252] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINALSolution 1:

[0253] For any set of parameters M0_Pi, a first solution could be to incrementally increase the parameter BV1 by a factor of 5 until the yield of each compound becomes constant, and then decrease the parameter BV4 until the yield of each compound becomes constant. This leads to the simulation of a setting associated with M0_Pi for which the flow of compounds through zones 1 and 4 has been eliminated, i.e., at the primary yield (hereafter referred to as "primary M0_Pi"). It is then possible to calculate, for the setting M0_Pi and for each compound, the yield difference 630 (or lost yield) by subtracting the effective yield (calculated without modifying the yield setting) from the primary yield, for example, using the following formula: Lost yield = Primary yield - Effective yield.

[0254] 15 Each resulting line can correspond to a training data point for training at least one model.

[0255] This performance difference is by definition zero for the "M0_Pi primary" setting.

[0256] The determination of the effective efficiency, calculated without modification of the 20 setting, is presented in figure 6 and repeated in the upper part of figure 7. The determination of the primary efficiency, for a BV1 max and a BV4 min, is presented in the lower part of figure 7 and the determination of the lost efficiency is presented in 630 in figure 7.

[0257] 25 Solution 2:

[0258] A second solution involves starting with a setting that prevents any compound from leaking through zones 1 and 4 ("primary M0_Pi"). The parameters BV1 and BV4 are then decreased and increased, respectively, until compound leakage through zones 1 and 4 occurs, with a limit value slightly above BV2 for the lowest BV1 value and a limit value slightly below BV3 for the highest BV4 value. Thus, a single primary M0_Pi setting will correspond to several M0_Pi settings.

[0259] 35. Using, for example, one of the two solutions presented above, the simulation of any parameter set M0_Pi makes it possible to generate data for each of the N compounds, i.e., N BV datasets, a

[0260] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL set of concentrations (profiles, extract and refine) and a lost yield value.

[0261] Construction / learning of a first artificial intelligence model 5 To enable the learning of this artificial intelligence model, a database was created from a simulation.

[0262] This simulation can be based on any of the published regulations, and from any of the existing software programs that are capable of calculating profile evolutions over time on several columns.

[0263] If we consider a MO simulation with N compounds, obtaining S11 from the dataset may involve using the simulation to vary the charge concentration parameters, the average mobile phase velocity, and the values ​​of each zone BV volume. Thus, for the MO simulation, obtaining S11 may involve running a large number of simulations varying one or another of the parameters, so as to generate "M0_P tot» simulations in total (these also include leak-free data from zone 1 or 4). This set of simulations then allows the generation of a database of N_M0 x M0_P tot 20 elements of BV, concentrations and lost yield.

[0264] In the invention, it is also necessary to use another simulation M1, generating M1_P tot simulation elements and for each as many data points as there are compounds of M1. Let N_M1 x M1_P tot elements.

[0265] This protocol is then repeated on other simulations M2 to M20, in order to obtain data on at least 21 simulations M0 to M20. With each repetition, each simulation generates as much information as there are compounds in the considered mixture N_Mt. This therefore allows us to obtain a set composed of

[0266]

[0267] 'M m _P tot data, each constituting a training sample.

[0268] Training can then be performed with all or part of this set, depending on the input data taken by the model. For example, training can be performed with the following input data:

[0269] - BV volume of raffinate,

[0270] 35 - BV volume of extract,

[0271] - 3 concentrations of the compound in question in each of zones 1 and 4, and

[0272] - the extracted and refined concentrations of the compound in question.

[0273] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL The output data is the lost yield of the compound considered. In Figure 8, for this model, each line 701 therefore corresponds to a training data point.

[0274] A major advantage of this artificial intelligence is that it is independent of the simulation and applies to every available data point.

[0275] In one example, the model's input parameters are the following twelve values ​​(Figure 10 shows the corresponding concentrations on the corresponding concentration profile): BVR, BVX, 3 concentrations of a compound in zone 4 (C4_p1, C4_p2, C4_p3), the concentration of the compound in the raffinate (T1), a concentration of the compound in zone 3 at the point closest to the end of the raffinate collection (C3_p1), a concentration of the compound in zone 2 at the point closest to the beginning of the extract (C2_p3), the concentration of the compound in the extract (T2), 3 concentrations of the compound in zone 1 (C1_p1, C1_p2, C1_p3).

[0276] 15 With:

[0277] - BVR: BV volume of raffinate,

[0278] - BVX: BV volume of extract,

[0279] - C4_pi: concentrations in zone 4 at the beginning, middle, and end of the loop,

[0280] 20 - T1: concentration of raffinate,

[0281] - C3_p1: first concentration after the raffinate,

[0282] - C2_p3: last concentration before the extract,

[0283] - T2: concentration of the extract, and

[0284] - C1_pi: concentrations in zone 1 at the beginning, middle and end of loop 25.

[0285] The output parameter is the lost yield.

[0286] The trained neural network has an architecture that may include:

[0287] - 1 to 3 layers (256 and 128 points),

[0288] - a ReLU function (acronym for "rectified linear unit") for hidden layers, and

[0289] - a linear function for the output layer.

[0290] The database: normalized data from databases 35 generated by simulation, for example for the following cases:

[0291] - lactic acid purification - Langmuir,

[0292] - purification of glucose-fructose sugars - anti-Langmuir, - purification of glycerol and salts - linear,

[0293] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL15 Langmuir, linear and anti-Langmuir isothermals, and 4 different speeds.

[0294] Learning outcome

[0295] 5 The following tables 1 and 2 illustrate in a non-exhaustive way the different construction options of the first AI model both on the structure of the input data for which the number of profile points of each zone has been modified, for which the volumes of fractions have or have not been used, and for different AI models and different associated parameters.

[0296] 10 In particular, Table 1 details the AI ​​parameters for each construction option (one row per option) and the results, while Table 2 presents the corresponding input data. In this example, two AI models were tested: the NN model and the Random Forest model. For the NN AI model, several layers (2, 3, or 4) were used. For the Random Forest model, several variants were tested: the Standard model, the Bootstrap model, and the Gradient Boosting model, which included variations on the number of trees and the depth of the random forest.

[0297] Table 1:

[0298]

[0299]

[0300] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL

[0301]

[0302] Table 2:

[0303]

[0304] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINALLa legend (last), (middle), (first) respectively designates concentration measurements taken at the end, middle and beginning of the area or sequence concerned.

[0305] 5

[0306] This demonstrates the flexibility of the approach and its compatibility with a large number of different AI models and settings.

[0307] The number of profiles, which constitute the training dataset for this first AI model, was 1,048,576. This is an order of magnitude; a smaller dataset, on the order of 200,000 profiles, can already yield interesting results. This set of profiles was obtained from different simulations involving a total of about forty compounds or families of compounds.

[0308] 15

[0309] Construction / training of a second artificial intelligence model To enable the training of this second artificial intelligence model, a database was created from a simulation.

[0310] This simulation can be based on a simulation as close as possible to the user's application, and from any of the existing software that is capable of calculating the yield to extract of the different compounds of the product to be purified.

[0311] If we consider a MO simulation with N compounds, obtaining S11 from the dataset may involve using simulation 25 to vary the charge concentration parameters, the average mobile phase velocity, and the values ​​of each zone BV volume. Thus, for the MO simulation, obtaining S11 may involve running a large number of conditions by varying one or another of the process parameters, thereby generating "M0_P tot » total simulation results (which also include leak-free data from zone 1 or 4). This set of simulation results then allows the generation of a database of N_M0 x M0_P tot elements of BV, concentrations and lost yield.

[0312] The learning can then be carried out with all or part of this set of 35 depending on the data taken as input by the model.

[0313] In one example, the model's input parameters are defined as follows. The number of input parameters may depend on

[0314] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL each model. In examples, if N is the number of compounds of interest in the mixture, N+2 parameters can be used. These N+2 parameters can be BVR, rX_i (i from 1 to N), and BVX with:

[0315] - BVR: BV volume of refinery.

[0316] 5 - rX_i: yields to extract of compounds i (i from 1 to N), ranked in order of their BV retention volume.

[0317] - BVX: BV volume of extract.

[0318] - Optionally, the concentration parameters of the product to be treated and the operating speeds.

[0319] The output parameter is the lost yield.

[0320] In figure 9, for this model, each set 702 therefore corresponds to a training data point.

[0321] The trained neural network has an architecture that can include 15:

[0322] - 3 layers (256, 128 and 64 points),

[0323] - a ReLU function (acronym for "rectified linear unit") for hidden layers, and

[0324] 20 - one or more linear function(s) for the output layer.

[0325] For each simulation element M0_Pi, the learning step S12 can consider as input data the set of yields of compounds 1 to N. The output data is the lost yield of a compound i.

[0326] 25 In this example, the S12 learning process may involve training the artificial intelligence model for each component of the mixture. In other words, the S12 learning process may involve determining a set of respective weights for each component. In the usage process, the S22 application of the function may involve determining the respective set of weights for the component in question, and applying the function with the determined set of weights.

[0327] Surprisingly, it has been observed that returns are a fundamental data point that allows us to predict lost returns.

[0328] 35 The addition as input data of the BVs of eluent, refiner and extract also constitutes an element of refining the accuracy of this artificial intelligence model.

[0329] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL For example, applying the second AI model to glucose-fructose separation and DPN family required a dataset of 28,800 data points to create one AI for glucose, one AI for fructose, and one AI for DPN families. More generally, the minimum number of data points is around 5,000 to 15,000, and can exceed 30,000.

[0330] Learning outcome

[0331] Tables 3 and 4 below illustrate the accuracy of the second AI model on an example of separating sucrose (dpn), glucose, and 10-fructose polymers. Specifically, Table 3 details the AI ​​parameters for each construction option (one row per option) and the corresponding input data, while Table 4 presents the results. In this example, two AI models were tested: the NN model and the Random Forest model. For the NN AI model, several layers (1, 2, or 3) were used. For the Random Forest model, several variants were tested: the Standard model, the Bootstrap model, and the Gradient Boosting model.

[0332] Table 3:

[0333]

[0334] 20

[0335] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINALTable 4:

[0336]

[0337] This demonstrates the flexibility of the approach and its compatibility with a large number of different AI models and parameterizations. It is important to note that, in this approach involving a second AI model, it is necessary to have either an AI model with a single output for each component of the simulation or a single AI model with as many outputs as there are components in the simulation.

[0338] Note that acceptable results are also obtained when the BVs of the extract and raffinate fractions are not used as input data.

[0339] Third AI model

[0340] Surprisingly, it was also discovered that the use of AI 15 can be applied to "duplicate" the behavior of a simulation. As explained previously, simulation is faster than experimental testing; however, simulating a set of process parameters can take anywhere from a few seconds to a few minutes. When generating the database for the second artificial intelligence model, we create a database that includes the process parameters (including the BVs applied to the process and, optionally, the concentration parameters of the product to be treated and the operating speeds), the effective yields, and the lost yields for each species. For the third AI model, it is possible to use machine learning based on the 25 process parameters as input parameters to predict the lost yields for each species as output parameters.

[0341] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL By extension, this third AI model can also include models from a learning based on the process parameters used as input to predict the primary yields of each of the species as output.

[0342] Once this third AI model is created, it can be used to predict the result of a simulation calculation: from a set of process parameters, primary, lost and effective yields are calculated, and this in less than a tenth of a second.

[0343] An example is given in Table 5 below for the simulation of fructose, glucose, and DPN family separation. Applying the third AI model to glucose, fructose, and DPN family separation required a dataset of 28,800 data points to create one AI for glucose, one for fructose, and one for the DPN families. More generally, the minimum number of data points is around 15,000, and can exceed 30,000.

[0344] Table 5:

[0345]

[0346] Results of using learned artificial intelligence models

[0347] The first two models described above operate with excellent accuracy when run on experimental input data. Such accuracy is also achieved when learning is based partially or exclusively on simulated data, regardless of the simulations chosen.

[0348] Figure 12 shows an example system, where the system is a client computer, for example, a user's workstation. The example client computer includes a central processing unit (CPU) 1010 connected to an internal communication bus 1000, and a random access memory (RAM) 1070 also connected to the bus. The client computer further has a graphics processing unit (GPU) 1110 which is associated with a video RAM 1100 connected to the bus. The video RAM 1100 is also known in the field as buffer memory. A mass storage device controller 1020 manages access to a mass storage device, such as a hard disk drive 1030.Mass storage devices suitable for the tangible embodiment of computer program instructions and data include all forms of non-volatile memory, including, for example, solid-state memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard drives and removable disks; and magneto-optical disks. All of the above-mentioned devices can be supplemented by, or incorporated into, specially designed application-specific integrated circuits (ASICs). A network adapter manages access to a network. The client computer may also include a haptic device, such as a cursor control device, a keyboard, or the like. A cursor control device is used in the client computer to allow the user to selectively position a cursor at any desired location on the screen.Furthermore, the cursor control device allows the user to select various commands (up to 25) and input control signals. The cursor control device includes several signal generation devices for inputting control signals into the system. Typically, a cursor control device is a mouse, whose button is used to generate the signals. Alternatively, or in addition, the client computer system may include a touchpad and / or a touchscreen.

[0349] The computer program may include instructions executable by a computer, the instructions including means to cause the aforementioned system to perform the processes. The program may be stored on any data storage medium, including system memory. The program may, for example, be implemented in digital electronic circuits, or in computer hardware, firmware, software, or combinations thereof. The program may be implemented in the form of a device, for example, a

[0350] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL product embodied tangibly in a machine-readable storage device to be executed by a programmable processor. The process steps can be carried out by a programmable processor executing a program of instructions to perform the functions of the 5 processes by operating on input data and generating results.

[0351] The processor can thus be programmable and coupled to receive data and instructions from a data storage system, at least one input device, and at least one output device. The application program can be implemented in a high-level procedural or object-oriented programming language, or in assembly or machine language if necessary. In all cases, the language can be compiled or interpreted. The program can be a full installation program or an update program. Applying the program to the system results in instructions to execute the processes. The computer program can be stored and executed on a server in a cloud computing environment, with the server communicating via a network with one or more clients.In this case, a processing unit executes the instructions contained in the program, thereby causing the execution of processes in the cloud-type computing environment.

[0352] Example of the application of artificial intelligence models

[0353] Consider the chromatographic separation of a sugar mixture comprising sugar polymers called dpn, glucose, and 25-fructose. During separation on a 4-column SSMB chromatographic system, each column 2 m long and 3 cm in diameter, applying the zone parameters below yielded an initial purification result:

[0354] BV1 = 0.668, BV2 = 0.58, BV3 = 0.63, BV4 = 0.53

[0355] When the system is stable, composition measurements are carried out in each zone in order to establish a profile and composition measurements of the average extracts and raffinates are carried out.

[0356] 35 The results obtained for the different compounds are presented in Tables 6 and 7 below (concentrations in g / L) and also in Figure 10:

[0357] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINALTable 6:

[0358]

[0359] Table 7:

[0360]

[0361] 5 The effective yields of the extracts of the species dpn, glucose and fructose were measured at 13.0%, 5.47% and 91.5% respectively, the BV of the raffinates and extracts are equal to 0.1 and 0.088 respectively.

[0362] Use of the first artificial intelligence

[0363] The first artificial intelligence was applied to each of the species: fructose, glucose and dpn with a model trained on the basis of the BV extract and raffinate, the concentration extract T2 and raffinate T1, and the compositions C4_p1, C4_pP2, C4_p3, C3_p1, C2_p3, C1_p1, C1_p2, C1_p3.

[0364] 15 For the fructose compound, the first artificial intelligence model applied to the data in the first row of the table above calculates a lost yield of 3.1%, corresponding to a zone 1 leakage.

[0365] For glucose, the second line of the table is used to calculate a lost yield equal to -1%, corresponding to a zone 4 leak.

[0366] 20 For the dpn family, the third line of the table is used to calculate a lost yield equal to -12.80%, corresponding to a zone 4 leak.

[0367] Figures 11.1, 11.2 and 11.3 schematically represent the distribution of the different yields in the chromatography system for fructose, glucose and the dpn family, and in particular the yield lost compared to the primary yield to the extract provided by the function, for an effective yield calculated according to Figure 10.

[0368] Use of the second family of artificial intelligence

[0369] 30 For the second family of artificial intelligence, three AI models were created for the species fructose, glucose, and dpn. Each of these

[0370] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL models uses BV raffinate and BV extract data as well as DPN, glucose and fructose yields.

[0371] The artificial intelligence model for fructose calculates a lost yield of 2.65% corresponding to a zone 1 leak.

[0372] 5 The artificial intelligence model for glucose calculates a lost yield of -0.7% corresponding to a leak in zone 4.

[0373] The artificial intelligence model for the dpn family calculates a lost yield of -13.0% corresponding to a zone 4 leak.

[0374] Summary of results and verification

[0375] The table below summarizes the results obtained and also calculates the primary yield for each of the compounds.

[0376]

[0377] 15 To verify the results obtained by the AI ​​models, it is necessary to make various additional adjustments.

[0378] Thus, for fructose, the expected yield if zone 1 losses are eliminated by increasing BV1 is 94.6% and 94.15%, according to the models. By increasing BV1 to 0.687 and then to 0.706, the effective fructose 20 yield increases to 93.5% and 93.6%, respectively. The plateau is reached, and therefore, we can conclude that there is no further yield loss, and thus, 93.6% is the primary yield. This means that the yield lost by zone 1 was 2.1%, while the AI ​​models estimated 3.1% and 2.65%.

[0379] If we now decide to lower zone 4 to verify the prediction of lost glucose yield by reducing zone 4 from 0.53 to 0.5, the effective glucose yield to the extract increases to 4.75%. This means that the yield lost by zone 4 was -0.72%, whereas the AI ​​models estimated -1% and -0.7%.

[0380] To perform the verification, the water volume had to be increased from 0.138 to 300.205. The yield loss assessments were accurate to within 1%, which is an acceptable error for making a decision to change the setting.

[0381] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL Consistency Criteria

[0382] As seen in the previous example, the AI ​​models according to the invention exhibit good performance; however, discrepancies remain between them, and the accuracy is on the order of 1%. The models' results are also dependent on the analytical quality, which is not always reliable.

[0383] "Hallucinations" can also occur with data that is, for example, far removed from the training data. In such circumstances, the applicant discovered a simple consistency criterion that can be implemented across all the AI ​​models used.

[0384] The consistency criterion consists of ordering the compounds in order of elution on a column and checking the growth of primary and lost yields.

[0385] In the previous example, if we rank the compounds in order of retention, we have: dpn family < Glucose < Fructose

[0386] 15

[0387] Using this order for the lost yield, we observe:

[0388] - -12.8% < -1% < 3.1% for the first model

[0389] - -13% < -0.7% < 2.65% for the second AI models

[0390] 20. Using this order for primary yield, we observe:

[0391] - 0.2 < 4.47% < 94.6% for the first model

[0392] - 0% < 4.77% < 94.15% for the second AI models

[0393] The primary yield must also be positive.

[0394] 25

[0395] When compounds have similar retentions, they should have similar lost and primary yields; therefore, a variation of 1 to 2% can be tolerated, or an average can be used for this product family before applying the yield growth verification criterion.

[0396] This criterion can be measured by counting the number of times the growth in lost and primary returns is not observed. This allows us to assess the consistency of the calculation result.

[0397] Naturally, this growth order based on the elution order corresponds to the standard used when considering the primary, effective, and 35% lost yields for the extract line. If the raffinate line is considered the reference, then the consistency criterion becomes a criterion of decrease relative to the elution orders.

[0398] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL Application to chromatographic processes with three or more fractions

[0399] There are processes that allow the collection of multiple fractions, not just an extract fraction and a raffinate fraction. These processes involve multi-column systems and the use of periodic or non-periodic subsequences, comprising steps of mobile phase injection, product injection, and collection of at least one extract and at least one raffinate fraction. Similar to two-fraction processes, they are considered in this application to fall within the family of simulated moving bed processes. Artificial intelligence models can also be used in these processes, depending on the specific method.

[0400] If the process sequences are periodic, meaning that the water and the product to be treated are injected several times in the cycle (as many times as there are columns, or half as many if the injection is performed on every other column), then the equivalents of zones 1 and 4 of the SMB process exist formally as in an SMB process, and profiling is performed in the same way. The yields lost through zones 1 and 4 are measurable with the same protocols, and the first artificial intelligence model is applicable.

[0401] The second and third artificial intelligence models apply 20 in all cases.

[0402] Whether the process is periodic or non-periodic, the second artificial intelligence model is applied by considering different input data for the model: instead of a single yield per species on a two-fraction process (for example, the extract yield, the raffinate yield being the complement to 100% of the extract yield), it is necessary to consider 2 yields on a 3-fraction process (the missing yield being the complement to 100% of the two yields considered), similarly it will be necessary to consider 3 yields on a 4-fraction process, etc.

[0403] Thus, for a simulation of a 3-fraction process, the base of 30 input data will consist of the effective yield of two of the three fractions, the volumes of the three fractions and always to calculate the yields lost by the equivalents of zones 1 and 4 of the process considered.

[0404] APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL

Claims

47 DEMANDS 1. A computer-implemented method for the automatic learning of a function configured to control a chromatographic separation process of a mixture comprising at least two compounds in a simulated moving-bed multicolumn system, the system comprising series-connected columns filled with stationary phase and mixing injection points, eluent injection points, extract collection points, and raffinate collection points cyclically switching between the columns, the system comprising zones 1, 2, 3, and 4, the function comprising at least one model configured to take as input operating data and to provide as output a lost yield relative to a primary yield for one of the compounds in the mixture, the learning method comprising: - obtaining (S11) a dataset comprising, for a set of training samples, 20 operating data points and associated lost efficiencies; and - machine learning (S12) of the function from the obtained dataset. 25 2. Learning method according to claim 1, wherein the primary yield is obtained by maximizing the volume of zone 1 and minimizing the volume of zone 4.

3. Learning method according to claim 1 or 2, wherein the operating data taken as input by at least one model include parameters including volumes of raffinate and extract.

4. A learning method according to any one of the 35 preceding claims, wherein the function comprises a first model, the operating data taken as input by the first model including measurements of APX.04540.DR1-2025.02.04-FirstFiling-APP-FINAL48 concentration at different nodes of the system, and / or concentration measurements in the extract and / or the raffinate.

5. A learning method according to claim 4, wherein 5 the concentration measurements at different nodes of the system comprise: - at least one concentration value taken in zone 4 and in zone 1; and - at least one concentration value taken in zone 2 and zone 3.

6. A training method according to claim 4 or 5, wherein the training samples include training samples for the following chromatographic separations 15: - several isotherms, for example of the Langmuir, linear and anti-Langmuir type, including but not limited to: o a purification of lactic acid, for example of the Langmuir type; 20 o a purification of glucose-fructose sugars, for example of the anti-Langmuir type; o a purification of glycerol and salts, for example of the linear type, the training samples can optionally include 25 training samples for different flow velocities in the system.

7. A learning method according to any one of the preceding claims, wherein the function comprises a second model, the operating data taken as input by the second model including yield measurements of all compounds from the chromatographic separation in the extract and / or the raffinate. 35 8. A learning method according to any one of the preceding claims, wherein the dataset is obtained by simulating the chromatographic separation process. APX.04540.DR1-2025.02.04-FirstFiling-APP-FINAL49 9. A learning method according to any one of the preceding claims, wherein the function comprises a third model, the operating data taken as input by the third model being process parameters. 5 10. A learning method according to any one of the preceding claims, wherein each model has an architecture comprising: - at least one hidden layer; - a rectified linear unit function (ReLU); and - an output layer including a linear function.

11. A method for using a function learned automatically according to the machine learning method of any one of the preceding claims, the method of use comprising: - obtaining (S21) operating data for a chromatographic separation process in a simulated moving bed multicolumn system; and 20 - an application (S22) of the function to the operating data obtained.

12. A method of use according to claim 11, wherein the method further comprises, after the application of the function, 25 a use (S23) of the lost yield supplied at the output by the function to control the chromatographic separation process.

13. A method of use according to claim 11 or 12, wherein the method further comprises: - implementation and / or use: o a first criterion of consistency of lost yield including a verification of the growth of lost yield for each 35 compound or family of compounds ordered according to their respective retention order, the first criterion being optionally measured by counting the number of times where the growth of yields APX.04540.DR1-2025.02.04-FirstFiling-APP-FINAL50 lost is inconsistent with the growth of retention orders; and / or o a second primary performance consistency criterion including a verification of the 5 growth of primary yield for each compound or family of compounds ordered according to their respective retention order, the second criterion being optionally measured by counting the number of times the growth of primary yields is inconsistent with the growth of retention orders.

14. Computer program comprising instructions for carrying out the machine learning method of one of claims 1 to 9 and / or the method of using one of claims 10 to 13 when said program is run on a computer.

15. Computer-readable storage medium on which the computer program of claim 14 is recorded.

16. System comprising a processor coupled to a memory, the memory having stored the computer program of claim 14. 25 APX.04540.DR1-2025.02.04-FirstFiling-APP -FINAL