Optimized arrangement method and device for one-dimensional acoustic diffuse reflection metasurface and storage medium

By combining MCMC and Q-Learning algorithms to optimize the acoustic diffuse reflection metasurface unit arrangement, the problem of global and local optimization imbalance in the prior art is solved, and the uniformity of the scattered sound field is improved.

CN120493425AActive Publication Date: 2025-08-15TONGLING UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510567416.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

When designing acoustic diffuse reflection metasurfaces, it is difficult to achieve a balance between global and local optimization, resulting in unsatisfactory diffuse reflection effects, and existing methods are prone to fall into local optimal solutions.

Method used

The MCMC algorithm is used for global search, combined with the Q-Learning algorithm for local fine adjustment, optimize the arrangement of metasurface units, and optimize the Q value table through the Metropolis-Hastings algorithm and Bellmann update formula to achieve uniformity of the scattered sound field.

Benefits of technology

The maximum uniformity of diffuse reflection scattering energy in all directions is achieved, local optimal trapping is avoided, and the uniformity of scattering sound field distribution is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493425A_ABST
    Figure CN120493425A_ABST
Patent Text Reader

Abstract

The invention discloses a one-dimensional acoustic diffuse reflection metasurface optimization arrangement method, which comprises the following steps of: L1, initializing metasurface units and corresponding unit arrangement, calculating sound pressure distribution of the metasurface units and acquiring a standard deviation of scattering sound pressure distribution; l2, global search of a unit arrangement scheme is achieved through MCMC annealing optimization, and the optimal arrangement scheme of the metasurface units is obtained to serve as the initial state of Q-Learning; l3, the arrangement is further optimized by using Q-Learning reinforcement learning, so that the metasurface scattering of the optimal unit arrangement scheme is more uniform, and the optimized metasurface arrangement is obtained; and L4, calculating the scattering sound pressure of the optimized metasurface, and drawing an optimized sound field distribution polar coordinate graph. According to the method, global search can be carried out at the initial stage of optimization, a better unit arrangement structure can be quickly converged, and local optimum is avoided; and local fine adjustment is carried out on the preliminary result in the later stage of optimization, so that the maximum uniformity of diffuse reflection scattering energy in each direction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of acoustic metamaterial technology, and in particular to a one-dimensional acoustic diffuse reflection metasurface optimization arrangement method, device, and storage medium that integrates MCMC and Q-Learning. Background Art

[0002] Acoustic metasurface is a new type of metamaterial that can achieve noise control, stealth technology and ultrasonic imaging. By designing a fine unit structure in the subwavelength scale, acoustic metasurface can achieve abnormal manipulation of sound waves, such as negative refraction, focusing, wave elimination and diffuse reflection. Among them, the diffuse reflection metasurface has the advantage of evenly dispersing the incident sound energy in all directions, which can effectively reduce the impact of sound energy concentration on noise control and echo suppression. However, to achieve efficient diffuse reflection effect, the key lies in the arrangement design of the metasurface unit structure.

[0003] In order to combine the various units in a reasonable way and thus make the scattered sound field uniformly distributed in space, existing technologies generally adopt traditional optimization methods such as genetic algorithms, simulated annealing algorithms, and particle swarm optimization. Although these methods can explore the solution space of metasurface arrangement to a certain extent, due to the discreteness and high nonlinearity of the arrangement problem itself, these methods are often prone to falling into local optimality. At the same time, a single global search strategy is difficult to take into account both global and local optimization requirements, resulting in unsatisfactory diffuse reflection effects.

[0004] In addition, as the Monte Carlo Markov Chain (MCMC) method has been introduced into the field of metasurface optimization due to its strong global search capability, it can quickly find a better initial arrangement through random perturbations and acceptance criteria. However, the MCMC algorithm has certain shortcomings in local fine optimization. On the other hand, although reinforcement learning (such as Q-Learning) can also be used in existing technologies as an adaptive learning method, local fine-tuning and fine-grained adjustment can be achieved by gradually learning the optimal strategy through interaction with the environment. However, the use of reinforcement learning alone is easily affected by the large dimension of the state space and insufficient training samples, and it is difficult to converge stably to the global optimal solution.

[0005] To this end, this application specifically proposes a one-dimensional acoustic diffuse reflection metasurface optimization arrangement method to solve the above technical problems. Summary of the Invention

[0006] The main purpose of the present invention is to provide a one-dimensional acoustic diffuse reflection metasurface optimization arrangement method. By first using the MCMC algorithm to perform a global search to obtain a preliminary optimal arrangement, and then using the Q-Learning algorithm to locally fine-tune the preliminary arrangement, the standard deviation of the scattered sound field distribution is further reduced, and the maximum uniformity of the diffuse reflection scattered energy in all directions is achieved, so as to solve the technical problems raised in the background technology.

[0007] The present invention adopts the following technical solutions to solve the above technical problems:

[0008] The optimized arrangement method of one-dimensional acoustic diffuse reflection metasurface includes:

[0009] L1. Initialize the metasurface unit and the corresponding unit arrangement, calculate its sound pressure distribution and obtain the standard deviation of the scattered sound pressure distribution;

[0010] L2. Use MCMC annealing optimization to perform a global search for unit arrangement schemes and obtain the optimal arrangement scheme of the metasurface units as the initial state of Q-Learning;

[0011] L3. Use Q-Learning reinforcement learning to further optimize the arrangement, making the metasurface scattering of the optimal unit arrangement more uniform and obtaining the optimized metasurface arrangement;

[0012] L4. Calculate the scattered sound pressure of the optimized metasurface and draw a polar coordinate diagram of the optimized sound field distribution.

[0013] Preferably, the specific operation process of the L1 step operation includes:

[0014] L11. Set N metasurface units, and each unit reflects from K predefined reflection angles {θ1,θ2,…,θ K}Select;

[0015] L12. Randomly initialize the metasurface unit arrangement S0 and calculate its sound pressure distribution. Set the objective function J as the standard deviation of the scattered sound pressure distribution, and we have:

[0016]

[0017] Among them, P i is the normalized sound pressure level at each angle, is the mean of the sound pressure at all angles, N is the number of discrete sampling points representing the scattering angle, and J represents the scattering uniformity error. The smaller its value, the more uniform the scattering.

[0018] Preferably, the specific operation process of the L2 step includes:

[0019] L21. Set the optimization hyperparameters, initial temperature T, annealing coefficient α, and maximum number of iterations MaxIter_MCMC;

[0020] L22. Using the Metropolis-Hastings algorithm, randomly select a cell i from 1 to N, randomly modify its reflection angle from K angles, and generate a new permutation S_new;

[0021] L23. After each iteration, reduce the temperature by calculating T = α * T and record the optimal solution;

[0022] L24. During the entire iterative process, continue to update the optimal arrangement S_best until T approaches 0 or reaches the maximum number of iterations MaxIter_MCMC.

[0023] Preferably, the specific operation process of generating the new arrangement S_new in step L22 includes:

[0024] a1. Calculate the change in the objective function ΔJ using the following formula:

[0025] ΔJ=J new -J old

[0026] Among them J new and J old are the variances of the metasurface scattered sound pressure before and after the change;

[0027] a2. Decide whether to accept S_new based on the Metropolis criteria:

[0028] If ΔE<0, that is, the new structural arrangement S_new makes the scattering more uniform, it is directly accepted; if ΔE≥0, the probability P=e-ΔJ / T is used to decide whether to accept S_new, where T is the annealing parameter.

[0029] Preferably, the specific operation process of the L3 step includes:

[0030] L31. Set the Q-value table Q(S,A) with initial settings of zero, learning rate μ, discount factor γ, exploration rate ε, and the number of Q-Learning training rounds MaxIter_Q;

[0031] L32. Select the current state S_t. In each iteration, for each unit in the current state S_t, the agent uses the ε-greedy strategy to select an action:

[0032] Randomly change the reflection angle of the specified cell from the current value to another candidate angle with probability ε;

[0033] Select the action with the highest Q value in the corresponding state in the Q table with probability 1-ε;

[0034] After the L33 selected action, the corresponding unit in the arrangement is adjusted to generate a new state S_t+1 and the corresponding optimal arrangement S_best.

[0035] Preferably, the specific operation process for generating S_best in the L33 step includes:

[0036] b1. Calculate and obtain the scattering uniformity J(S_t+1) of the current arrangement S_t+1, and then calculate the reward R_t. The calculation formula is:

[0037] R_t = J_best - J(S_t+1)

[0038] where J_best is the optimal metasurface scattered sound pressure;

[0039] If J(S_t+1) < J_best, then R_t is positive, that is, the uniformity of the new state is improved, encouraging optimization;

[0040] b2. Update the Q value using the Bellman update formula: Update the optimal solution S_best. If J(S_t+1) < J_best, then S_best = S_t+1, and continue to iterate using the following formula:

[0041] Q(S t , A t ) = Q(S t , A t ) + μ·(R t + γmax a' Q(S t+1 , a') - Q(S t , A t ))

[0042] where (S t , A t ) is a set of state-action pairs, Q(S t , A t ) is the current Q value of the state-action pair (S t , A t ); R t is the immediate reward of the current action; γ is the discount factor, and 0 ≤ γ ≤ 1, which is used to control the importance of future rewards, and μ is the learning rate, which is used to determine the update amplitude of the Q value; max a' Q(S t+1 , a') represents taking the maximum value of the Q values corresponding to all possible actions a' in the next state S_t+1; a' represents any one of all possible actions in the next state S_t+1; R t + γmax a' Q(S t+1,a') is the target value, R t +γmax a' Q(S t+1 ,a')-Q(S t ,A t ) is the temporal difference error, which represents the deviation between the current estimate and the “target”;

[0043] Select the maximum Q value obtained after the optimal action as the updated Q value;

[0044] b3. Repeat the Q-value update operation until the predetermined number of rounds is reached or the Q-table change is less than the preset value. The training ends and the system outputs the final optimized hypersurface arrangement S_best.

[0045] In another aspect, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.

[0046] On the other hand, the present invention further discloses a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.

[0047] As can be seen from the above technical solution, the present invention provides a method for optimizing the arrangement of a one-dimensional acoustic diffuse reflection metasurface. Compared with the prior art, the present invention has the following advantages:

[0048] 1. By setting up a Markov chain Monte Carlo (MCMC) algorithm in the early stage of optimization, the present invention can play a global search role, thereby quickly converging to a more optimal unit arrangement structure and avoiding falling into a local optimum. By introducing the Q-Learning reinforcement learning algorithm in the later stage of optimization, the preliminary results are locally fine-tuned to further reduce the standard deviation of the scattered sound field distribution and achieve maximum uniformity of diffuse scattered energy in all directions.

[0049] 2. The present invention combines the scattering uniformity of the current arrangement with the reward calculation during the iterative selection of the unit arrangement structure state using the Q-value table, and iteratively updates the Q-value. This enables stable convergence to the global optimal solution during the reinforcement learning process, facilitating the output of the final optimized metasurface arrangement.

[0050] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become easy to understand through the following description. Of course, it is not necessary to achieve all of the above-mentioned advantages simultaneously in order to implement any product of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0052] Figure 1 Schematic diagram of the overall method flow of the present invention;

[0053] Figure 2 Schematic diagram of the acoustic pressure distribution of the initial arrangement metasurface of the present invention;

[0054] Figure 3 Schematic diagram of the acoustic pressure distribution of the metasurface after MCMC optimization of the present invention;

[0055] Figure 4 Schematic diagram of the acoustic pressure distribution of the metasurface after MCMC and Q-Learning optimization of the present invention. DETAILED DESCRIPTION

[0056] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. In the absence of conflict, the embodiments in this application and the features in the embodiments can be combined with each other. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0057] In the embodiment, see Figures 1 to 4 .

[0058] like Figure 1 The one-dimensional acoustic diffuse reflection metasurface optimization arrangement method proposed in the embodiment of the present invention includes the following steps:

[0059] Step 1: Start (initialization):

[0060] (1) Set N metasurface units;

[0061] (2) Each unit is reflected from K predefined reflection angles {θ1,θ2,…,θ K}Select;

[0062] (3) Randomly initialize the metasurface unit arrangement S0;

[0063] (4) Calculate its sound pressure distribution and set the objective function J as the standard deviation of the scattered sound pressure distribution:

[0064]

[0065] Where: P iis the normalized sound pressure level (dB) at each angle, is the mean of the sound pressure at all angles, N is the number of discrete sampling points representing the scattering angle, and J represents the scattering uniformity error; a smaller value indicates more uniform scattering. Set the optimization hyperparameters: initial temperature T, annealing coefficient α, and maximum number of iterations MaxIter_MCMC.

[0066] At this time, the supersurface acoustic pressure distribution is as follows Figure 2 shown.

[0067] Step 2: MCMC annealing optimization (global search):

[0068] Its main goal is to explore different unit arrangement schemes and find the optimal solution S_best as the initial state of Q-Learning.

[0069] (5) Using the Metropolis-Hastings algorithm, randomly select a unit i (from 1 to N), modify its reflection angle (randomly selected from K angles), and generate a new arrangement S_new.

[0070] (a) Calculate the change in the objective function ΔJ, which is: ΔJ = J new -J old ;

[0071] Among them J new and J old are the variances of the metasurface scattered sound pressure before and after the change, respectively.

[0072] (b) Decide whether to accept S_new based on the Metropolis criteria:

[0073] If ΔE<0 (the new solution is better), that is, the new structural arrangement S_new makes the scattering more uniform, it is directly accepted.

[0074] If ΔE≥0, the probability P=e-ΔJ / T (T is the annealing parameter) is used to decide whether to accept S_new.

[0075] (6) After each iteration, reduce the temperature T: T = α * T; record the optimal solution, and continuously update the optimal arrangement S_best during the entire iteration process until T approaches 0 or reaches the maximum number of iterations MaxIter_MCMC.

[0076] At this time, the supersurface acoustic pressure distribution is as follows Figure 3 shown.

[0077] Step 3: Q-Learning reinforcement learning (local optimization):

[0078] Objective: Based on S_best, use reinforcement learning to further optimize the arrangement to make the scattering more uniform.

[0079] (7) Set up the Q-value table Q(S,A) (initialized to zero), the learning rate μ (controlling the Q-value update speed), the discount factor γ (measuring the importance of future rewards), the exploration rate ε (used to control the balance between random exploration and greedy selection), and the number of Q-Learning training rounds MaxIter_Q.

[0080] (8) Select the current state S_t (initialized from the MCMC result S_best); in each iteration round, for each cell in the current state S_t, the agent uses the ε-greedy policy to select an action: with probability ε, randomly change the reflection angle of a certain cell from the current value to another candidate angle (the set of its actions is all possible reflection angles (there are K angles here)). With probability 1 - ε, select the action with the highest Q-value in the corresponding state in the Q-table, that is, select the action that maximizes the expected reward.

[0081] (9) After selecting the action, adjust the corresponding cell in the permutation to generate a new state S_t+1.

[0082] (a) Calculate the reward R_t: Calculate the scattering uniformity J(S_t+1) of the current permutation S_t+1; calculate the reward R_t = J_best - J(S_t+1). If J(S_t+1) < J_best, then R_t is positive, that is, the uniformity of the new state is improved (the standard deviation is reduced), encouraging optimization.

[0083] (b) Update the Q-value using the Bellman update formula: Update the optimal solution S_best. If J(S_t+1) < J_best, then S_best = S_t+1, and continue to iterate using the following formula:

[0084] Q(S t ,A t ) = Q(S t ,A t ) + μ·(R t + γmax a' Q(S t+1 ,a') - Q(S t ,A t ))

[0085] where (S t ,A t ) is a state-action pair, Q(S t ,A t ) is the current Q-value of the state-action pair (S t ,A t ); R tis the immediate reward for the current action; γ is the discount factor, and 0≤γ≤1, which is used to control the importance of future rewards; μ is the learning rate, which is used to determine the Q value update amplitude; max a' Q(S t+1 ,a') means that in the next state S_t+1, the Q value corresponding to all possible actions a' is the maximum; a' represents any one of all possible actions in the next state S_t+1; R t +γmax a' Q(S t+1 ,a') is the target value, R t +γmax a' Q(S t+1 ,a')-Q(S t ,A t ) is the temporal difference error, which represents the deviation between the current estimate and the “target”;

[0086] The maximum Q value obtained after selecting the optimal action is used as the updated Q value.

[0087] (c) After completing one Q-value update, proceed to the next step. Repeat until the predetermined number of rounds is reached or the Q-table changes are very small. The training ends and the system outputs the final optimized hypersurface arrangement S_best.

[0088] Step 4: Calculate the optimized scattered sound pressure:

[0089] (10) Output the optimal arrangement S_best and calculate the scattering characteristics of the optimized metasurface.

[0090] (11) Draw the polar coordinate diagram of the optimized metasurface acoustic field distribution, such as Figure 4 shown.

[0091] In summary, the algorithm adopted in this application is divided into a two-stage optimization strategy:

[0092] Phase 1: Using the Markov Chain Monte Carlo (MCMC) method and the Metropolis-Hastings rule, a global search is performed based on the initial random arrangement to quickly converge to a better unit arrangement structure and avoid falling into the local optimum.

[0093] Phase II: Based on the MCMC results, Q-Learning reinforcement learning is used to further optimize the arrangement and locally fine-tune the preliminary results to further reduce the standard deviation of the scattered sound field distribution and achieve maximum uniformity of the diffuse scattered energy in all directions.

[0094] In a specific embodiment, in order to optimize the 1D metasurface arrangement, the following setting parameters are substituted into the specific operation steps of the above embodiment:

[0095] Setting parameters:

[0096] The metasurface is set to contain N = 40 units; the predefined reflection angles, k = 5, are 15°, 30°, 45°, 60°, and 80° respectively;

[0097] Set the initial temperature T = 1.0, the annealing coefficient α = 0.95; the number of MCMC iterations MaxIter_MCMC = 800;

[0098] Set Q-Learning parameters: learning rate α = 0.05, discount factor γ = 0.9, exploration rate ε = 0.3. Q-Learning iteration number MaxIter_Q = 2000;

[0099] The optimization process at this time is as follows: initially randomly permutate S0, run MCMC sampling to obtain the preliminary optimized permutation S_init; run Q-Learning reinforcement learning to obtain the final optimization result S_best;

[0100] The final experimental results are:

[0101] The standard deviation of the initial random arrangement is 5.69, and the corresponding sound pressure distribution is as follows: Figure 2 As shown; the standard deviation after MCMC optimization is 4.11, and the corresponding sound pressure distribution is as follows Figure 3 As shown; after MCMC and 2000 rounds of training, the optimized standard deviation is 2.44, and the corresponding sound pressure distribution is as follows Figure 4 As shown in the figure, the uniformity of the metasurface is improved by 57.12% compared with the initial random arrangement and by 40.63% compared with the MCMC optimization.

[0102] In another aspect, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.

[0103] On the other hand, the present invention further discloses a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.

[0104] In another embodiment provided in the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any one-dimensional acoustic diffuse reflection metasurface optimization arrangement method in the above-mentioned embodiments.

[0105] It is understandable that the system provided by the embodiment of the present invention corresponds to the method provided by the embodiment of the present invention, and the explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts of the above method.

[0106] The embodiment of the present application further provides an electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus.

[0107] Memory for storing computer programs;

[0108] The processor is used to implement the above-mentioned one-dimensional acoustic diffuse reflection metasurface optimization arrangement method when executing the program stored in the memory.

[0109] The communication bus mentioned in the above electronic device can be a peripheral component interconnect standard bus or an extended industry standard architecture bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0110] The communication interface is used for communication between the above electronic device and other devices.

[0111] The memory may include a random access memory, or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0112] The above-mentioned processor can be a general-purpose processor, including a central processing unit, a network processor, etc.; it can also be a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component.

[0113] It should also be noted that electronic devices also include terminal devices, which can also be referred to as terminals, user equipment, mobile stations, mobile terminals, etc. Terminal devices can be mobile phones, smart TVs, wearable devices, tablet computers, computers with wireless transceiver functions, virtual reality terminal devices, augmented reality terminal devices, wireless terminals in industrial control, wireless terminals in unmanned driving, wireless terminals in remote surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, etc. The embodiments of this application do not limit the specific technologies and specific device forms used by terminal devices.

[0114] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium, or a semiconductor medium (e.g., a solid-state hard disk).

[0115] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

[0116] In addition, it should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture. If the specific posture changes, the directional indications will also change accordingly.

[0117] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the meaning of "and / or" appearing throughout the text includes three parallel schemes. Taking "A and / or B" as an example, it includes scheme A, or scheme B, or schemes in which A and B are satisfied at the same time. In addition, in the embodiments of the present invention, "multiple" refers to more than two. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the ability of ordinary technicians in this field to implement. When the combination of technical solutions is mutually contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

Claims

1. One-dimensional acoustic diffuse reflection metasurface optimization arrangement method, characterized in that: include: L1. Initialize the metasurface unit and the corresponding unit arrangement, calculate its sound pressure distribution and obtain the standard deviation of the scattered sound pressure distribution; L2. Use MCMC annealing optimization to achieve a global search for unit arrangement schemes and obtain the optimal arrangement scheme of the metasurface units as the initial state of Q-Learning; L3. Use Q-Learning reinforcement learning to further optimize the arrangement, making the metasurface scattering of the optimal unit arrangement more uniform and obtaining the optimized metasurface arrangement; L4. Calculate the scattered sound pressure of the optimized metasurface and draw a polar coordinate diagram of the optimized sound field distribution.

2. The one-dimensional acoustic diffuse reflection metasurface optimization arrangement method according to claim 1, characterized in that: The specific operation process of the L1 step operation includes: L11. Set N metasurface units, and each unit reflects from K predefined reflection angles {θ1,θ2,…,θ K }Select; L12. Randomly initialize the metasurface unit arrangement S0 and calculate its sound pressure distribution. Set the objective function J as the standard deviation of the scattered sound pressure distribution, and we have: Among them, P i is the normalized sound pressure level at each angle, is the mean of the sound pressure at all angles, N is the number of discrete sampling points representing the scattering angle, and J represents the scattering uniformity error. The smaller its value, the more uniform the scattering.

3. The one-dimensional acoustic diffuse reflection metasurface optimization arrangement method according to claim 2, characterized in that: The specific operation process of the L2 step includes: L21. Set the optimization hyperparameters, initial temperature T, annealing coefficient α, and maximum number of iterations MaxIter_MCMC; L22. Using the Metropolis-Hastings algorithm, randomly select a cell i from 1 to N, randomly modify its reflection angle from K angles, and generate a new permutation S_new; L23. After each iteration, reduce the temperature by calculating T = α * T and record the optimal solution; L24. During the entire iterative process, continue to update the optimal arrangement S_best until T approaches 0 or reaches the maximum number of iterations MaxIter_MCMC.

4. The one-dimensional acoustic diffuse reflection metasurface optimization arrangement method according to claim 3, characterized in that: The specific operation process of generating the new arrangement S_new in step L22 includes: a1. Calculate the change in the objective function ΔJ using the following formula: ΔJ=J new -J old Among them J new and J old are the variances of the metasurface scattered sound pressure before and after the change; a2. Decide whether to accept S_new based on the Metropolis criteria: If ΔE<0, that is, the new structural arrangement S_new makes the scattering more uniform, it is directly accepted; if ΔE≥0, the probability P=e-ΔJ / T is used to decide whether to accept S_new, where T is the annealing parameter.

5. The one-dimensional acoustic diffuse reflection metasurface optimization arrangement method according to claim 3, characterized in that: The specific operation process of the L3 step includes: L31. Set the Q-value table Q(S,A) with initial settings of zero, learning rate μ, discount factor γ, exploration rate ε, and the number of Q-Learning training rounds MaxIter_Q; L32. Select the current state S_t. In each iteration, for each unit in the current state S_t, the agent uses the ε-greedy strategy to select an action: Randomly change the reflection angle of the specified cell from the current value to another candidate angle with probability ε; Select the action with the highest Q value in the corresponding state in the Q table with probability 1-ε; L33. After selecting the action, adjust the corresponding units in the arrangement to generate a new state S_t+1 and the corresponding optimal arrangement S_best.

6. The one-dimensional acoustic diffuse reflection metasurface optimization arrangement method according to claim 5, characterized in that: The specific operation process of generating S_best in the L33 step includes: b1. Calculate the scattering uniformity J(S_t+1) of the current arrangement S_t+1, and then calculate the reward R_t. The calculation formula is: R_t=J_best-J(S_t+1) Wherein, J_best is the optimal metasurface scattering sound pressure; If J(S_t+1) < J_best, then R_t is positive, that is, the uniformity of the new state is improved, encouraging optimization; b2. Update the Q value: Update the optimal solution S_best. If J(S_t+1) < J_best, then S_best = S_t+1, and continue to iterate using the following formula: Q(S t ,A t )=Q(S t ,A t )+μ·(R t +γmax a, Q(S t+1 ,a’)-Q(S t ,A t )) Among them, (S t ,A t ) is a set of state-action pairs, Q(S t ,A t ) is a state-action pair (S t ,A t )’s current Q value; R t is the immediate reward for the current action; γ is the discount factor, and 0≤γ≤1, which is used to control the importance of future rewards; μ is the learning rate, which is used to determine the Q value update amplitude; max a, Q(S t+1 ,a') means that in the next state S_t+1, the Q value corresponding to all possible actions a' is the maximum; a' represents any one of all possible actions in the next state S_t+1; R t +γmax a, Q(S t+1 ,a') is the target value, R t +γmax a, Q(S t+1 ,a')-Q(S t ,A t ) is the temporal difference error, which represents the deviation between the current estimate and the "target"; Select the maximum Q value obtained after choosing the optimal action as the updated Q value; b3. Repeat the Q value update operation until the training ends when the predetermined number of rounds is reached or the change in the Q table is less than the preset value, and the system outputs the finally optimized metasurface arrangement S_best.

7. A computer-readable storage medium, characterized in that A computer program is stored. When the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 6.

8. A computer device, characterized in that: It includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Online adjustment method for control parameters of autonomous aircraft based on MCMC optimized Q learning

    CN107885086A

  • Five-mode metasurface underwater local sound field intensity regulation and control method based on deep learning

    CN115688524A

  • Local sound field regulation and control method based on user-defined loss function and multi-feature constraint

    CN116631369A

  • Phase holographic imaging device based on acoustic metamaterial and design method

    CN119864003A

  • Approximated objective function for monte carlo algorithm

    US20230306290A1