Hand disc tuning method, system and device and medium
Through reinforcement learning model assisting tuning equipment for hand disc tuning, the problem of tuning in the prior art relying on professional manual labor is solved, and an efficient and low-cost tuning process is achieved.
Patent Information
- Application Number
- CN202510223913.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-27
AI Technical Summary
In the prior art, the tuning of hand discs depends on professional tuners, which is complex and time-consuming, has high labor costs and low efficiency.
By obtaining the model information of the hand disc and the sound area position, determining the target vibration frequency, collecting deformation amount and vibration frequency, inputting it into the reinforcement learning model, predicting the action value, instructing the tuning device to perform tuning operations, and updating the model according to the reward value.
It reduces the labor cost of hand disc tuning, improves tuning efficiency, facilitates mass production and calibration, and improves user experience.
Smart Images

Figure CN120148445A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a tuning method, system, device, and storage medium for a handpan. Background Art
[0002] The handpan, also known as the Hang Drum or UFO Drum, is a modern percussion instrument. It consists of two hemispherical metal shells, usually made of nitride steel, and the metal surface is hardened through a special manufacturing process. The shape of the handpan is similar to an upside-down flying saucer, so it is also called the "UFO Drum". The sound of the handpan is soft and distant, with a strong resonance effect, and the timbre is rich and variable, which is loved by people.
[0003] As a precision percussion instrument, the pitch and timbre of the handpan are crucial for the performance effect. During the production or use process, it is necessary to tune the handpan to keep it in good pitch and timbre and improve the playing experience of the handpan. However, in related technologies, the tuning of the handpan often relies on professional tuners, the implementation process is complex and time-consuming, the labor cost of tuning is high, and the efficiency is low.
[0004] In summary, the problems existing in the related technologies need to be solved urgently. Summary of the Invention
[0005] An object of this application is to solve at least to some extent one of the technical problems existing in the related technologies.
[0006] To this end, an object of an embodiment of this application is to provide a tuning method, system, device, and storage medium for a handpan.
[0007] To achieve the above technical object, the technical solutions adopted in the embodiments of this application include:
[0008] On the one hand, an embodiment of this application provides a tuning method for a handpan, and the method includes:
[0009] Obtain the model information and the sound area position of the target handpan, and determine the corresponding target vibration frequency according to the model information and the sound area position;
[0010] Collect the first deformation amount at the sound area position of the target handpan, and strike the sound area position with a predetermined knocking force, and collect the first vibration frequency corresponding to the sound area position;
[0011] Input the first deformation amount and the first vibration frequency as state values into a reinforcement learning model, and predict and output a first action value through the reinforcement learning model; wherein, the first action value is used to indicate the pressing position and pressing distance of the pressing head of the tuning device on the target handpan;
[0012] According to the first motion value, perform a tuning operation on the target handpan. After the tuning operation, collect the second deformation amount at the sound region position of the target handpan, and strike the sound region position with a predetermined striking force, and collect the second vibration frequency corresponding to the sound region position;
[0013] According to the second vibration frequency and the target vibration frequency, determine the reward value corresponding to the tuning operation of the reinforcement learning model in the current round;
[0014] According to the reward values in multiple rounds of tuning operations, update the parameters of the reinforcement learning model to obtain a trained reinforcement learning model;
[0015] Based on the trained reinforcement learning model, tune the target handpan through a tuning device.
[0016] In addition, according to the handpan tuning method of the above embodiments of the present application, the following additional technical features may also be included:
[0017] Further, in an embodiment of the present application, the tuning device includes an infrared ranging mechanism; the collecting of the first deformation amount at the sound region position of the target handpan includes:
[0018] Move the infrared ranging mechanism directly above the sound region position;
[0019] Detect the distance from the infrared ranging mechanism to the sound region position;
[0020] According to the distance, determine the first deformation amount at the sound region position.
[0021] Further, in an embodiment of the present application, the determining of the reward value corresponding to the tuning operation of the reinforcement learning model in the current round according to the second vibration frequency and the target vibration frequency includes:
[0022] Calculate the absolute value of the difference between the second vibration frequency and the target vibration frequency;
[0023] According to the absolute value, determine the reward value corresponding to the tuning operation of the reinforcement learning model in the current round;
[0024] Wherein, the size of the absolute value is negatively correlated with the size of the reward value.
[0025] Further, in an embodiment of the present application, the determining of the reward value corresponding to the tuning operation of the reinforcement learning model in the current round according to the absolute value includes:
[0026] Obtain a preset linear coefficient and exponential coefficient;
[0027] Calculate a first value by using the absolute value as the base number and the exponential coefficient as the exponent;
[0028] Calculate a second value according to the product of the linear coefficient and the first value;
[0029] Take the opposite of the second value to obtain the reward value.
[0030] Further, in an embodiment of the present application, the updating the parameters of the reinforcement learning model according to the reward value in multiple rounds of tuning operations includes:
[0031] Obtain a preset maximum number of rounds per episode;
[0032] If the number of execution rounds of the tuning operation reaches the maximum number of rounds per episode in the current training episode, obtain all the reward values in the current training episode;
[0033] Sum all the reward values in the current training episode to obtain a total reward value;
[0034] Update the parameters of the reinforcement learning model according to the total reward value.
[0035] Further, in an embodiment of the present application, the tuning the target handpan by a tuning device based on the trained reinforcement learning model includes:
[0036] Collect a third deformation amount at the sound zone position of the target handpan, strike the sound zone position with a predetermined striking force, and collect a third vibration frequency corresponding to the sound zone position;
[0037] Input the third deformation amount and the third vibration frequency as state values into the trained reinforcement learning model, and predict and output a target action value through the trained reinforcement learning model; wherein, the target action value is used to indicate the pressing position and pressing distance of the pressing head of the tuning device on the target handpan;
[0038] Perform a tuning operation on the target handpan by the tuning device according to the target action value.
[0039] On the other hand, an embodiment of the present application provides a handpan tuning system, the system includes:
[0040] An acquisition unit, configured to acquire model information and sound zone position of a target handpan, and determine a corresponding target vibration frequency according to the model information and the sound zone position;
[0041] The acquisition unit is configured to acquire a first deformation amount at the sound area position of the target handpan, and strike the sound area position with a predetermined knocking force to acquire a first vibration frequency corresponding to the sound area position.
[0042] The prediction unit is configured to input the first deformation amount and the first vibration frequency as state values into a reinforcement learning model, and predict and output a first action value through the reinforcement learning model; wherein, the first action value is used to indicate the pressing position and pressing distance of the indenter of the tuning device on the target handpan.
[0043] The operation unit is configured to perform a tuning operation on the target handpan according to the first action value, acquire a second deformation amount at the sound area position of the target handpan after the tuning operation, and strike the sound area position with a predetermined knocking force to acquire a second vibration frequency corresponding to the sound area position.
[0044] The evaluation unit is configured to determine a reward value corresponding to the reinforcement learning model in the current round of tuning operation according to the second vibration frequency and the target vibration frequency.
[0045] The update unit is configured to update the parameters of the reinforcement learning model according to the reward values in multiple rounds of tuning operations to obtain a trained reinforcement learning model.
[0046] The execution unit is configured to tune the target handpan through a tuning device based on the trained reinforcement learning model.
[0047] On the other hand, an embodiment of the present application provides an electronic device, including:
[0048] At least one processor;
[0049] At least one memory for storing at least one program;
[0050] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned tuning method for the handpan.
[0051] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to implement the above-mentioned tuning method for the handpan when executed by the processor.
[0052] The advantages and beneficial effects of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application:
[0053] The tuning method, system, device and storage medium of a handpan disclosed in the embodiments of the present application. The method obtains the model information and the pitch area position of the target handpan, determines the corresponding target vibration frequency according to the model information and the pitch area position; collects the first deformation amount at the pitch area position of the target handpan, and strikes the pitch area position with a predetermined striking force, and collects the first vibration frequency corresponding to the pitch area position; inputs the first deformation amount and the first vibration frequency as state values into a reinforcement learning model, and predicts and outputs a first action value through the reinforcement learning model; wherein, the first action value is used to indicate the pressing position and pressing distance of the indenter of the tuning device on the target handpan; according to the first action value, perform a tuning operation on the target handpan, after the tuning operation, collect the second deformation amount at the pitch area position of the target handpan, and strike the pitch area position with a predetermined striking force, and collect the second vibration frequency corresponding to the pitch area position; determine the reward value corresponding to the tuning operation of the reinforcement learning model in the current round according to the second vibration frequency and the target vibration frequency; update the parameters of the reinforcement learning model according to the reward values in multiple rounds of tuning operations to obtain a trained reinforcement learning model; based on the trained reinforcement learning model, tune the target handpan through a tuning device. This method can effectively reduce the labor cost of handpan tuning, improve the tuning efficiency, facilitate batch production and calibration of handpans, and is beneficial to improving the user experience by training a reinforcement learning model to assist in the handpan tuning task. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the accompanying drawings of the relevant technical solutions in the embodiments of the present application or the prior art. It should be understood that the accompanying drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0055] Figure 1 Schematic diagram of the implementation environment of a handpan tuning method provided in the embodiments of the present application;
[0056] Figure 2 Schematic flowchart of a handpan tuning method provided in the embodiments of the present application;
[0057] Figure 3 Schematic diagram of the structure of a handpan tuning system provided in the embodiments of the present application;
[0058] Figure 4 Schematic diagram of the structure of an electronic device provided in the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] The present application will be further described below in conjunction with the accompanying drawings of the specification and specific embodiments. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0060] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0062] 1) Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0063] 2) Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning (deep learning) usually includes technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0064] 3) Reinforcement Learning (RL), an important branch of machine learning, focuses on how an agent learns the optimal behavioral strategy through interaction with the environment. In reinforcement learning, the agent continuously tries different actions and adjusts its behavior according to the reward or punishment signals obtained from the environment to maximize the long-term cumulative reward.
[0065] The handpan, also known as the Hang Drum or UFO Drum, is a modern percussion instrument. It consists of two hemispherical metal shells, usually made of NITRIDE steel, and the metal surface is hardened through a special manufacturing process. The shape of the handpan resembles an upside-down flying saucer, so it is also called the "UFO Drum". The sound of the handpan is soft and distant, with a strong resonance effect, and the timbre is rich and variable, which is loved by people.
[0066] As a precision percussion instrument, the pitch and timbre of the handpan are crucial for the performance effect. During the production or use process, it is necessary to tune the handpan to keep it in good pitch and timbre and improve the playing experience of the handpan. However, in the related technology, the tuning of the handpan often relies on professional tuners, the implementation process is complex and time-consuming, the labor cost of tuning is high, and the efficiency is low.
[0067] In view of this, an embodiment of the present application provides a tuning method for a handpan. The method obtains the model information and the sound area position of the target handpan, determines the corresponding target vibration frequency according to the model information and the sound area position; collects the first deformation amount at the sound area position of the target handpan, and strikes the sound area position with a predetermined striking force, and collects the first vibration frequency corresponding to the sound area position; inputs the first deformation amount and the first vibration frequency as state values into a reinforcement learning model, and predicts and outputs a first action value through the reinforcement learning model; wherein, the first action value is used to indicate the pressing position and pressing distance of the indenter of the tuning device on the target handpan; according to the first action value, perform a tuning operation on the target handpan, collect the second deformation amount at the sound area position of the target handpan after the tuning operation, and strike the sound area position with a predetermined striking force, and collect the second vibration frequency corresponding to the sound area position; determine the reward value corresponding to the tuning operation of the reinforcement learning model in the current round according to the second vibration frequency and the target vibration frequency; update the parameters of the reinforcement learning model according to the reward values in multiple rounds of tuning operations to obtain a trained reinforcement learning model; based on the trained reinforcement learning model, tune the target handpan through a tuning device. This method can effectively reduce the labor cost of handpan tuning, improve the tuning efficiency, facilitate batch production and calibration of handpans, and is conducive to improving the user experience by training a reinforcement learning model to assist in the handpan tuning task.
[0068] Please refer to Figure 1 , Figure 1 FIG. shows a schematic diagram of an implementation environment of a tuning method for a handpan provided in an embodiment of the present application. In this implementation environment, the main software and hardware entities involved include a terminal device 110 and a background server 120. The terminal device 110 and the background server 120 are communicatively connected.
[0069] Specifically, the tuning method for a handpan provided in an embodiment of the present application can be executed independently on the terminal device 110 side, or can be executed independently on the background server 120 side, or can be executed based on data interaction between the terminal device 110 and the background server 120.
[0070] Among them, the terminal device 110 in the above embodiments may include a mobile phone, a computer, a smart wearable device, a PDA device, a smart voice interaction device, a vehicle-mounted terminal, etc., but is not limited thereto. The background server 120 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0071] A communication connection can be established between the terminal device 110 and the background server 120 through a wireless network or a wired network. The wireless network or wired network uses standard communication technologies and / or protocols. The network can be set to the Internet or any other network, such as including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network.
[0072] Of course, it can be understood that Figure 1 the implementation environment in Figure 1 is only some optional application scenarios of the handpan tuning method provided in the embodiments of the present application. The actual application is not fixed to the
[0073] software and hardware environment shown.
[0074] Next, in combination with the introduction of the foregoing implementation environment, a handpan tuning method provided in the embodiments of the present application will be introduced and described.
[0074] Please refer to Figure 2 Figure 2 which is a schematic diagram of a handpan tuning method provided in the embodiments of the present application. The handpan tuning method includes but is not limited to:
[0075] Step 210: Obtain the model information and the sound area position of the target handpan, and determine the corresponding target vibration frequency according to the model information and the sound area position;
[0076] Step 220: Collect the first deformation amount at the sound area position of the target handpan, and strike the sound area position with a predetermined knocking force, and collect the first vibration frequency corresponding to the sound area position;
[0077] Step 230: Input the first deformation amount and the first vibration frequency as state values into the reinforcement learning model, and predict and output a first action value through the reinforcement learning model; wherein, the first action value is used to indicate the pressing position and pressing distance of the indenter of the tuning device on the target handpan.
[0078] Step 240: According to the first action value, perform a tuning operation on the target handpan. After the tuning operation, collect the second deformation amount at the sound zone position of the target handpan, and strike the sound zone position with a predetermined knocking force, and collect the second vibration frequency corresponding to the sound zone position.
[0079] Step 250: Determine the reward value corresponding to the reinforcement learning model in the current round of tuning operation according to the second vibration frequency and the target vibration frequency.
[0080] Step 260: Update the parameters of the reinforcement learning model according to the reward values in multiple rounds of tuning operations to obtain a trained reinforcement learning model.
[0081] Step 270: Based on the trained reinforcement learning model, tune the target handpan through the tuning device.
[0082] In the embodiment of the present application, a tuning method for a handpan is provided. This method can effectively reduce the labor cost of handpan tuning, improve the tuning efficiency, facilitate batch production and calibration of handpans, and is conducive to improving the user experience by training a reinforcement learning model to assist in the handpan tuning task.
[0083] The method in the embodiment of the present application uses reinforcement learning technology. In the application of reinforcement learning, concepts such as agent, environment, state, action, and reward are included. Among them, an agent is an object that perceives changes in the environment and executes corresponding actions, which can be relevant control components or software programs, etc.; the environment is the world where the agent interacts. For the application scenario of the present application, the environment is the usage environment of the target handpan; when the agent and the environment interact, the agent will observe the current state of the environment and decide what action to take and execute according to the current state. After the agent executes the action, the environment will change (it may also remain unchanged), and then a new state will be generated. And the agent will receive a feedback reward from the environment, and this reward represents the impact of the agent's action on the environment, such as whether the environment becomes better or worse relative to the pre-set target. The goal of the agent is to maximize the cumulative reward obtained from the environment, and reinforcement learning is the way for the agent to achieve this goal through learning behavior.
[0084] Specifically, in the embodiments of the present application, the handpan to be tuned is denoted as the target handpan. In the embodiments of the present application, the model information of the target handpan and the position of the pitch area to be tuned can be obtained. According to the model information and the pitch area position of the target handpan, the vibration frequency corresponding to the target handpan when working in this pitch area under normal circumstances can be determined, and it is denoted as the target vibration frequency. In the embodiments of the present application, there are no restrictions on the specific model, pitch area position, and magnitude of the target vibration frequency of the target handpan.
[0085] For the target handpan, relevant data involved in the tuning process can be trained based on the reinforcement learning algorithm. Specifically, each time a tuning operation is performed, the deformation amount at the pitch area position of the target handpan before the tuning operation is performed can be collected. In the embodiments of the present application, it is denoted as the first deformation amount; then, the pitch area position can be struck with a predetermined striking force, and the vibration frequency corresponding to the pitch area position is collected and denoted as the first vibration frequency. In the embodiments of the present application, the first deformation amount and the first vibration frequency can be used as state values and input into the reinforcement learning model, and an action value is predicted and output by the reinforcement learning model, denoted as the first action value. Here, the first action value can be used to indicate the pressing position and pressing distance of the indenter of the tuning device on the target handpan.
[0086] After determining the first action value, the tuning operation can be performed on the target handpan according to the first action value. Then, after the tuning operation is performed, the deformation amount at the pitch area position of the target handpan can be continuously collected. In the embodiments of the present application, it is denoted as the second deformation amount. Then, the pitch area position can be struck with the same predetermined striking force as before, and the corresponding vibration frequency is collected and denoted as the second vibration frequency. In the embodiments of the present application, the reward value corresponding to the tuning operation of the reinforcement learning model in the current round can be determined according to the second vibration frequency and the target vibration frequency. The above process of performing the tuning operation can be executed multiple times, and each time it is executed, a reward value will be obtained. In the embodiments of the present application, the parameters of the reinforcement learning model can be updated according to the reward values in multiple rounds of tuning operations, so as to obtain a trained reinforcement learning model. There are no restrictions on the specific number of iterative rounds in the present application, and it can be flexibly set according to actual needs.
[0087] After obtaining the trained reinforcement learning model, the target handpan can be tuned using the tuning device based on the reinforcement learning model. By outputting the action value that the tuning device needs to execute through the reinforcement learning model, the tuning of the target handpan can be conveniently achieved. It can be understood that the method in the embodiments of the present application does not need to rely on professional tuners to complete the tuning task of the handpan, which can reduce the labor cost of tuning and improve the tuning efficiency.
[0088] Specifically, in some embodiments, the tuning device includes an infrared ranging mechanism; collecting the first deformation amount at the sound region position of the target handpan includes:
[0089] Moving the infrared ranging mechanism directly above the sound region position;
[0090] Detecting the distance to the sound region position through the infrared ranging mechanism;
[0091] Determining the first deformation amount at the sound region position according to the distance.
[0092] In the embodiments of the present application, an infrared ranging mechanism may be provided in the tuning device. When measuring the deformation amount at the sound region position of the target handpan, the infrared ranging mechanism can be moved directly above the sound region position. Through the infrared ranging mechanism, the distance to the sound region position can be detected. According to the measured distance, the deformation amount at the sound region position can be determined, such as the first deformation amount and the second deformation amount, etc. Exemplarily, for example, in the embodiments of the present application, when starting to execute the task, an initial distance can be measured, and the deformation amount corresponding to the initial distance is recorded as 0. After obtaining new distance data subsequently, it can be compared with the initial distance to determine the corresponding deformation amount.
[0093] Specifically, determining the reward value corresponding to the tuning operation of the reinforcement learning model in the current round according to the second vibration frequency and the target vibration frequency includes:
[0094] Calculating the absolute value of the difference between the second vibration frequency and the target vibration frequency;
[0095] Determining the reward value corresponding to the tuning operation of the reinforcement learning model in the current round according to the absolute value;
[0096] Wherein, the size of the absolute value is negatively correlated with the size of the reward value.
[0097] In the embodiments of the present application, when determining the reward value corresponding to each round in the training process, the difference between the second vibration frequency and the target vibration frequency can be calculated, and the absolute value of the difference is taken. Then, according to the absolute value, the reward value corresponding to the tuning operation in the current round is determined. It is easy to understand that the closer the second vibration frequency is to the target vibration frequency, the closer the tuning operation in the current round is to the target requirement, that is, the better the effect. At this time, the reward value can be determined as a larger value; on the contrary, the more the second vibration frequency is not close to the target vibration frequency, the more the tuning operation in the current round deviates from the target requirement, that is, the worse the effect. At this time, the reward value can be determined as a smaller value. In other words, the size of the obtained absolute value can be negatively correlated with the size of the reward value. The specific relationship between the two is not limited in the present application.
[0098] For example, in some embodiments, determining the reward value corresponding to the tuning operation of the reinforcement learning model in the current round according to the absolute value includes:
[0099] Obtain a preset linear coefficient and an exponential coefficient;
[0100] Using the absolute value as the base and the exponential coefficient as the exponent, calculate a first value;
[0101] Calculate a second value according to the product of the linear coefficient and the first value;
[0102] Take the opposite of the second value to obtain the reward value.
[0103] In the embodiments of the present application, when determining the reward value corresponding to the tuning operation, a preset linear coefficient and an exponential coefficient can be obtained, and then, the reward value can be calculated through the following formula:
[0104] Reward=-k×|ob[0]-target_Hz|^p p
[0105] In the formula, Reward represents the reward value, k represents the linear coefficient, ob[0] represents the second vibration frequency, target_Hz represents the target vibration frequency, and p represents the exponential coefficient.
[0106] Specifically, in some embodiments, updating the parameters of the reinforcement learning model according to the reward values in multiple rounds of tuning operations includes:
[0107] Obtain a preset maximum number of rounds per episode;
[0108] If the number of execution rounds of the tuning operation reaches the maximum number of rounds per episode in the current training episode, obtain all the reward values in the current training episode;
[0109] Sum all the reward values in the current training episode to obtain the total reward value;
[0110] Update the parameters of the reinforcement learning model according to the total reward value.
[0111] In the embodiments of the present application, when updating the parameters of the reinforcement learning model, the entire training process can be divided into several rounds, and each round can include several turns. In the embodiments of the present application, the turns corresponding to each round are denoted as the maximum number of turns in the round. For each training round, the turn number can be incremented by 1 after each tuning operation is executed. If the execution turn number of the tuning operation reaches the maximum number of turns in the current training round, the sum of all reward values within this training round can be obtained to get a comprehensive reward value, which is denoted as the total reward value in the embodiments of the present application. After obtaining the total reward value, the parameters of the reinforcement learning model can be updated based on it, and then this training round can be combined.
[0112] It should be noted that in the embodiments of the present application, if the predetermined tuning requirement is met after a certain tuning operation is executed (for example, the second vibration frequency is relatively close to the target vibration frequency), the training of this round can also be ended, and the present application does not limit this.
[0113] Specifically, in some embodiments, tuning the target handpan by the tuning device based on the trained reinforcement learning model includes:
[0114] Collect the third deformation amount at the sound area position of the target handpan, and strike the sound area position with a predetermined striking force, and collect the third vibration frequency corresponding to the sound area position;
[0115] Input the third deformation amount and the third vibration frequency as state values into the trained reinforcement learning model, and predict and output a target action value through the trained reinforcement learning model; wherein, the target action value is used to indicate the pressing position and pressing distance of the indenter of the tuning device on the target handpan;
[0116] Perform a tuning operation on the target handpan by the tuning device according to the target action value.
[0117] In the embodiments of the present application, when using the tuning device to tune the target handpan, the deformation amount at the sound area position of the target handpan can be collected and denoted as the third deformation amount. Then, the sound area position can be struck with a predetermined striking force, and the vibration frequency corresponding to the sound area position can be collected and denoted as the third vibration frequency. Then, the third deformation amount and the third vibration frequency can be input as state values into the trained reinforcement learning model to predict and output an action value, which is denoted as the target action value in the embodiments of the present application. According to the target action value, the target handpan can be tuned.
[0118] Generally speaking, the action execution rules for tuning the target handpan can be summarized as follows: when the current vibration frequency (i.e., the third vibration frequency) is higher than the target vibration frequency, the action execution mechanism of the tuning device uses the upper pressing head to press down on the handpan; when the current vibration frequency is lower than the target vibration frequency, the action execution mechanism of the tuning device uses the lower pressing head to press down on the handpan. Of course, the specific situation depends on the actual needs, and this application does not limit it.
[0119] Referring to Figure 3 , this application embodiment also provides a handpan tuning system, including:
[0120] An acquisition unit 310, configured to acquire the model information and the pitch area position of the target handpan, and determine the corresponding target vibration frequency according to the model information and the pitch area position;
[0121] A collection unit 320, configured to collect the first deformation amount at the pitch area position of the target handpan, and strike the pitch area position with a predetermined knocking force, and collect the first vibration frequency corresponding to the pitch area position;
[0122] A prediction unit 330, configured to input the first deformation amount and the first vibration frequency as state values into a reinforcement learning model, and predict and output a first action value through the reinforcement learning model; wherein, the first action value is used to indicate the pressing position and pressing distance of the pressing head of the tuning device on the target handpan;
[0123] An operation unit 340, configured to perform a tuning operation on the target handpan according to the first action value, collect the second deformation amount at the pitch area position of the target handpan after the tuning operation, and strike the pitch area position with a predetermined knocking force, and collect the second vibration frequency corresponding to the pitch area position;
[0124] An evaluation unit 350, configured to determine the reward value corresponding to the reinforcement learning model in the current round of tuning operation according to the second vibration frequency and the target vibration frequency;
[0125] An update unit 360, configured to update the parameters of the reinforcement learning model according to the reward values in multiple rounds of tuning operations, and obtain a trained reinforcement learning model;
[0126] An execution unit 370, configured to tune the target handpan through a tuning device based on the trained reinforcement learning model.
[0127] It can be understood that the content in the above method embodiments is applicable to this system embodiment. The functions specifically implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0128] Referring to Figure 4 , an embodiment of the present application provides an electronic device, including:
[0129] At least one processor 410;
[0130] At least one memory 420, configured to store at least one program;
[0131] When the at least one program is executed by the at least one processor 410, the at least one processor 410 implements the above-mentioned tuning method of the handpan.
[0132] Similarly, the content in the above method embodiments is applicable to the embodiments of this electronic device. The functions specifically implemented by the embodiments of this electronic device are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0133] The embodiment of the present application further provides a computer-readable storage medium, in which a program executable by the processor 410 is stored. The program executable by the processor 410 is used to execute the above-mentioned tuning method of the handpan when executed by the processor 410.
[0134] Similarly, the content in the above method embodiments is applicable to the embodiments of this computer-readable storage medium. The functions specifically implemented by the embodiments of this computer-readable storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0135] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present application are provided by way of example for a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.
[0136] In addition, although the present application has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present application. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present application as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.
[0137] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0138] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0139] More specific examples (nonexhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0140] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0141] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0142] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the claims and their equivalents.
[0143] The above has specifically described the preferred embodiments of the present application, but the present application is not limited to the embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
Claims
1. A handpan tuning method, characterized in that: The method comprises: Acquire the model information and the sound zone position of the target handpan, and determine the corresponding target vibration frequency according to the model information and the sound zone position; collecting a first deformation amount at the sound zone position of the target handpan, and striking the sound zone position with a predetermined striking force to collect a first vibration frequency corresponding to the sound zone position; Inputting the first deformation amount and the first vibration frequency as state values into a reinforcement learning model, and predicting and outputting a first action value through the reinforcement learning model; wherein the first action value is used to indicate the pressing position and pressing distance of the pressure head of the tuning device on the target handpan; According to the first action value, a tuning operation is performed on the target handpan, and after the tuning operation, a second deformation amount of the target handpan at the sound zone position is collected, and the sound zone position is struck with a predetermined striking force to collect a second vibration frequency corresponding to the sound zone position; Determining a reward value corresponding to the reinforcement learning model in a current round of tuning operation according to the second vibration frequency and the target vibration frequency; According to the reward values in multiple rounds of tuning operations, parameters of the reinforcement learning model are updated to obtain a trained reinforcement learning model; Based on the trained reinforcement learning model, the target handpan is tuned by a tuning device.
2. A handpan tuning method according to claim 1, characterized in that: The tuning device includes an infrared distance measuring mechanism; the collecting of the first deformation amount at the sound zone position of the target handpan includes: Move the infrared distance measuring mechanism to the position just above the sound zone; Detecting the distance from the sound zone position by the infrared distance measuring mechanism; A first deformation amount at the sound zone position is determined according to the distance.
3. The handpan tuning method according to claim 1, characterized in that: The step of determining, according to the second vibration frequency and the target vibration frequency, a reward value corresponding to the reinforcement learning model in the current round of tuning operation includes: calculating an absolute value of a difference between the second vibration frequency and the target vibration frequency; Determining, according to the absolute value, a reward value corresponding to the reinforcement learning model in the current round of tuning operation; The absolute value is negatively correlated with the reward value.
4. The handpan tuning method according to claim 3, characterized in that: Determining, according to the absolute value, a reward value corresponding to the reinforcement learning model in the current round of tuning operation includes: Get the preset linear coefficients and exponential coefficients; The absolute value is used as a base and the exponential coefficient is used as an exponent to calculate a first value; Calculate a second value according to the product of the linear coefficient and the first value; The reward value is obtained by taking the opposite of the second value.
5. The handpan tuning method according to claim 1, characterized in that: The step of updating parameters of the reinforcement learning model according to the reward values in the multiple rounds of tuning operations includes: Get the preset maximum number of rounds; If the execution round number of the tuning operation in the current training round reaches the maximum round number of the round, all reward values in the current training round are obtained; Sum all reward values in the current training round to get the total reward value; According to the total reward value, parameters of the reinforcement learning model are updated.
6. A handpan tuning method according to any one of claims 1-5, characterized in that: The step of tuning the target handpan by a tuning device based on the trained reinforcement learning model includes: collecting a third deformation amount at the sound zone position of the target handpan, and striking the sound zone position with a predetermined striking force to collect a third vibration frequency corresponding to the sound zone position; Inputting the third deformation amount and the third vibration frequency as state values into the trained reinforcement learning model, and predicting and outputting a target action value through the trained reinforcement learning model; wherein the target action value is used to indicate the pressing position and pressing distance of the pressure head of the tuning device on the target handpan; According to the target action value, a tuning operation is performed on the target handpan through the tuning device.
7. A tuning system for a handpan, characterized in that: The system comprises: An acquisition unit, used to acquire model information and a tone zone position of a target handpan, and determine a corresponding target vibration frequency according to the model information and the tone zone position; A collecting unit, used for collecting a first deformation amount at the sound zone position of the target handpan, and striking the sound zone position with a predetermined striking force to collect a first vibration frequency corresponding to the sound zone position; a prediction unit, configured to input the first deformation amount and the first vibration frequency as state values into a reinforcement learning model, and predict and output a first action value through the reinforcement learning model; wherein the first action value is used to indicate a pressing position and a pressing distance of a pressing head of a tuning device on the target handpan; an operating unit, configured to perform a tuning operation on the target handpan according to the first action value, collect a second deformation amount at the sound zone position of the target handpan after the tuning operation, and strike the sound zone position with a predetermined striking force to collect a second vibration frequency corresponding to the sound zone position; an evaluation unit, configured to determine a reward value corresponding to the reinforcement learning model in a current round of tuning operation according to the second vibration frequency and the target vibration frequency; An updating unit, configured to update parameters of the reinforcement learning model according to the reward values in multiple rounds of tuning operations to obtain a trained reinforcement learning model; An execution unit is used to tune the target handpan through a tuning device based on the trained reinforcement learning model.
8. The handpan tuning system according to claim 7, characterized in that: The updating unit is specifically used for: Get the preset maximum number of rounds; If the execution round number of the tuning operation in the current training round reaches the maximum round number of the round, all reward values in the current training round are obtained; Sum all reward values in the current training round to get the total reward value; According to the total reward value, parameters of the reinforcement learning model are updated.
9. An electronic device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a handpan tuning method according to any one of claims 1 to 6.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to implement a handpan tuning method according to any one of claims 1 to 6 when executed by the processor.
Citation Information
Patent Citations
RBF neural network-based various musical instruments tone tuning method
CN107705775A
Tuning of a drum
CN108292496A
Piano tuning method and system based on neural network
CN109991842A
Piano remote tone tuning and diagnosing method and system based on deep learning
CN110444226A
Automatic tuning terminal equipment based on artificial intelligence
CN112634874A