Audio signal processing device and method for replacing music using machine learning model

KR1020260123977APending Publication Date: 2026-08-14GAUDI AUDIO LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
KR1020260023010
Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-07
Filing Date
2026-02-04
Publication Date
2026-08-14

Smart Images

  • Figure P1020260023010_ABST
    Figure P1020260023010_ABST
Patent Text Reader

Abstract

An audio signal processing device for replacing music using a machine learning model is disclosed. The audio signal processing device includes a processor. The processor receives an input music audio signal, obtains a replacement target segment including a signal containing music to be replaced from the input music audio signal, determines a final candidate target segment among the plurality of candidate segments based on the similarity between the plurality of candidate segments obtained by the machine learning model and the replacement target segment, and outputs an audio signal in which the replacement target segment in the input music audio signal is replaced with the final candidate target segment.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to an audio signal processing method and apparatus for replacing music using a machine learning model. Background Technology

[0002] With the advancement of communication technology and the popularization of streaming services, a single piece of content is being consumed across multiple countries. Copyright laws and agreements vary by nation. Therefore, it is necessary to verify whether music included in the content is being used without consultation with the copyright holder. In particular, music copyright issues can arise when attempting to distribute past content—which was originally produced without global distribution in mind—through global services.

[0003] Accordingly, it is necessary to replace music included in content that poses copyright issues in the regions where the content is distributed with music that does not. However, manually selecting and replacing the problematic music can be excessively time-consuming and costly. Therefore, a music signal processing device capable of performing this task efficiently is required. In particular, a music signal processing device utilizing machine learning models that determine the characteristics of music is necessary. The problem to be solved

[0004] The embodiment of the present invention aims to provide an audio signal processing method and apparatus for replacing music using a machine learning model. means of solving the problem

[0005] According to one embodiment of the present invention, an audio signal processing device for replacing music using a machine learning model includes a processor. The processor receives an input music audio signal, obtains a replacement target segment including a signal containing music to be replaced from the input music audio signal, determines a final candidate target segment among the plurality of candidate segments based on the similarity between the plurality of candidate segments obtained by the machine learning model and the replacement target segment, and outputs an audio signal in which the replacement target segment in the input music audio signal is replaced with the final candidate target segment.

[0006] The similarity between the above plurality of candidate segments and the above replacement target segment may include similarity in musical characteristics.

[0007] The similarity between the plurality of candidate segments and the replacement target segment may include the similarity between the sound pressure envelope contour of the plurality of candidate segments and the sound pressure envelope contour of the replacement target segment.

[0008] Each of the above-mentioned plurality of candidate segments is a portion of music corresponding to each of the above-mentioned plurality of candidate segments, and the audio signal processing device may store information indicating a start time and an end time indicating the music corresponding to each of the above-mentioned plurality of candidate segments and the portion of music corresponding to each of the above-mentioned plurality of candidate segments.

[0009] The processor can adjust the loudness of the final candidate target segment when replacing the replacement target segment with the final candidate target segment in the input music audio signal.

[0010] When the processor replaces the replacement target segment with the final candidate target segment in the input music audio signal, it can adjust the loudness of the final candidate segment based on the spectrum or instantaneous loudness level of the replacement target segment.

[0011] When the processor replaces the replacement target segment with the final candidate target segment in the input music audio signal, if the length of the final candidate segment is shorter than the length of the replacement target segment, a part of the final candidate segment may be appended before or after the final candidate segment.

[0012] The processor can divide a song into a plurality of replacement target segments and independently determine the final candidate segment for each of the plurality of replacement target segments.

[0013] When the processor obtains the replacement target segment from the input audio signal, if the difference between the end time of the first replacement target segment and the start time of the second replacement target segment is less than or equal to a reference value, it can merge them into one replacement target segment.

[0014] According to an embodiment of the present invention, a method of operation of an audio signal processing device for replacing music using a machine learning model comprises: receiving an input music audio signal; obtaining a replacement target segment including a signal containing music to be replaced from the input music audio signal; determining a final candidate target segment among a plurality of candidate segments based on the similarity between a plurality of candidate segments obtained by the machine learning model and the replacement target segment; and outputting an audio signal in which the replacement target segment in the input music audio signal is replaced with the final candidate target segment.

[0015] The similarity between the above plurality of candidate segments and the above replacement target segment may include similarity in musical characteristics.

[0016] The similarity between the plurality of candidate segments and the replacement target segment may include the similarity between the sound pressure envelope contour of the plurality of candidate segments and the sound pressure envelope contour of the replacement target segment.

[0017] Each of the above-mentioned plurality of candidate segments is a portion of music corresponding to each of the above-mentioned plurality of candidate segments, and the audio signal processing device may store information indicating a start time and an end time indicating the music corresponding to each of the above-mentioned plurality of candidate segments and the portion of music corresponding to each of the above-mentioned plurality of candidate segments.

[0018] The above method of operation may further include the step of adjusting the loudness of the final candidate target segment when replacing the replacement target segment with the final candidate target segment in the input music audio signal.

[0019] The step of adjusting the loudness of the final candidate target segment may include adjusting the loudness of the final candidate segment based on the spectrum or instantaneous loudness level of the replacement target segment when replacing the replacement target segment with the final candidate target segment in the input music audio signal.

[0020] The above operation method may further include the step of appending a part of the final candidate segment to the front or back of the final candidate segment when the length of the final candidate segment is shorter than the length of the replacement target segment when the replacement target segment is replaced in the input music audio signal.

[0021] The step of determining the final candidate target segment among the plurality of candidate segments may further include the step of dividing one song into a plurality of replacement target segments and independently determining the final candidate segment for each of the plurality of replacement target segments.

[0022] The step of acquiring the above-mentioned replacement target segment may include, when acquiring the above-mentioned replacement target segment from the input audio signal, merging into one replacement target segment if the difference between the end time of the first replacement target segment and the start time of the second replacement target segment is less than or equal to a reference value. Effects of the invention

[0023] The apparatus and method according to an embodiment of the present invention can provide an audio signal processing method and apparatus for replacing music using a machine learning model. Brief explanation of the drawing

[0024] FIG. 1 shows a block diagram of an audio signal processing device according to an embodiment of the present invention. FIG. 2 shows a block diagram of an operation in which an audio signal processing device according to an embodiment of the present invention replaces music. FIG. 3 shows the operation of an audio signal processing device according to an embodiment of the present invention analyzing the shape of candidate music. FIG. 4 shows an algorithm used by an audio signal processing device according to an embodiment of the present invention to analyze the form of candidate music. FIGS. 5 and 6 show specific operations of an audio signal processing device generating candidate segments. FIGS. 7 to 9 show a machine learning algorithm used by an audio signal processing device according to an embodiment of the present invention to acquire musical characteristics from candidate segments and replacement target segments. FIG. 10 shows an audio signal processing device according to an embodiment of the present invention generating a replacement target segment from input music audio and generating embedding information including musical characteristics of the replacement target segment. Specific details for implementing the invention

[0025] Embodiments of the present invention are described below with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present invention in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification have been given similar reference numerals. Additionally, when a part is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0026] An audio signal processing device according to an embodiment of the present invention includes at least one processor. The operation of audio signal processing described in the present invention may be the operation of an instruction set operating in the processor included in the audio signal processing device.

[0027] FIG. 1 shows a block diagram of an audio signal processing device according to an embodiment of the present invention.

[0028] The audio signal processing device determines the music segment among the candidate music segments in the database that has the most similar characteristics to the characteristics of the input music audio signal. The audio signal processing device replaces the input music audio signal with the music segment having the most similar characteristics. At this time, the multiple candidate music segments stored in the database may consist of multiple pieces of music each segmented into multiple sections. Additionally, the audio signal processing device may include a database containing information representing the characteristics of each music segment among the candidate music segments in the database.

[0029] FIG. 2 shows a block diagram of an operation in which an audio signal processing device according to an embodiment of the present invention replaces music.

[0030] An audio signal processing device can classify the form of each of the multiple segments included in the candidate music. The audio signal processing device can generate candidate music segmentation based on the forms of the classified multiple segments. At this time, the audio signal processing device can distinguish the multiple segments based on the characteristics of the multiple segments included in the candidate music. At this time, the audio signal device can generate one or more consecutive segments as a single candidate segment. The multiple segments may correspond to at least one of an intro, pre-chorus, chorus, verse, bridge, or coda. This is explained in detail through FIGS. 3 and 4. In addition, a method for generating a candidate segment by combining one or more segments is explained through FIGS. 5 and 6.

[0031] The audio signal processing device can store embeddings indicating the features of each generated candidate segment. To this end, the audio signal device can extract features from each generated music segment. A method for the audio signal processing device to extract the features of a music segment is described through FIGS. 7 to 9.

[0032] The audio signal processing device can generate a database of candidate segments and embeddings representing the characteristics of the candidate segments through the operation described above. According to a specific embodiment, the audio signal processing device may use a generated database without directly generating a database of candidate segments and embeddings representing the characteristics of the candidate segments.

[0033] The audio signal processing device can divide the audio signal to be replaced into multiple segments. In this case, the operation of the audio signal processing device generating multiple segments may be the same as the operation of generating candidate segments. Additionally, the audio signal processing device extracts the characteristics of the generated segments. The audio signal processing device compares the characteristics of the segments with the characteristics of the candidate segments and replaces the segment to be replaced with a candidate segment based on the comparison result. In this case, the audio signal processing device may replace the segment to be replaced with a candidate segment that has characteristics most similar to those of the segment to be replaced.

[0034] When the audio signal processing device replaces the segment to be replaced with a candidate segment, the audio signal processing device can adjust the loudness of the candidate segment. Specifically, the audio signal processing device can adjust the loudness of the candidate segment using time enveloping. This will be explained again after the description of Fig. 10.

[0035] FIG. 3 shows the operation of an audio signal processing device according to an embodiment of the present invention analyzing the shape of candidate music. FIG. 4 shows an algorithm used by an audio signal processing device according to an embodiment of the present invention to analyze the shape of candidate music.

[0036] An audio signal processing device can classify the form of each of the multiple sections included in the candidate music. At this time, the audio signal processing device can distinguish the multiple sections according to the characteristics of the multiple sections included in the candidate music. The characteristics of the sections may be compositional characteristics according to the development of the song. For example, the multiple sections may correspond to at least one of an intro, pre-chorus, chorus, verse, bridge, or coda.

[0037] In addition, the audio signal processing device can demix audio signals for each source, such as instruments and vocals, and process the spectrogram of the demixed signal using a transformer module to distinguish multiple sections according to the characteristics of the sections.

[0038] FIGS. 5 and 6 show specific operations of an audio signal processing device generating candidate segments.

[0039] An audio signal processing device can generate candidate segments by combining multiple sections of candidate music. In this case, the audio signal processing device can generate candidate segments by combining only sections that are consecutive in the development sequence of the candidate music. Additionally, the audio signal processing device can include any one section in multiple segments. In the embodiment of FIG. 6, the audio signal processing device generates 15 candidate segments from candidate music distinguished by five sections: intro, coda, crus, coda, and bus. In this case, each section is included in multiple candidate segments, and a candidate segment includes one or more consecutive sections in the development sequence of the candidate music. Through this, the audio signal processing device can generate segments with various characteristics from a small number of candidate musics.

[0040] In this embodiment, the audio signal processing device may not store candidate segments in the form of individual files. Specifically, the audio signal processing device may store information indicating time intervals corresponding to candidate music and segments generated from the candidate music. For example, the information indicating time intervals corresponding to segments in the candidate music may include the segment start time and end time. Through this, the audio signal processing device can store information regarding a relatively large number of segments in a small storage space.

[0041] FIGS. 7 to 9 show a machine learning algorithm used by an audio signal processing device according to an embodiment of the present invention to acquire musical characteristics from candidate segments and replacement target segments.

[0042] An audio signal processing device can obtain musical characteristics from candidate segments and replacement target segments using the music tagging transformer model illustrated in FIG. 7 (minz Won, et al., "Semi-supervised Music Tagging Transformer", International Society for Music Information Retrieval (ISMIR) 2021). The musical characteristics may include at least one of genre, instrumentation, tempo, or key. The music tagging transformer is a model that attaches tagging information, such as the genre or mood of the music, and the encoder of the model provides an interpretation of the genre and mood typically expressed in the tagging.

[0043] In addition, the audio signal processing device can obtain musical characteristics from candidate segments and replacement target segments using the text-to-music representation model illustrated in Fig. 8 (Seungheon Doh, et al., "Toward Universal Text-to-Music Retrieval", IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2023). This model is a model that searches for music corresponding to general text, and generates semantic-based clustering where music is represented as text by learning the Joint Embedding Space of text and music.

[0044] In addition, the audio signal processing device can obtain musical characteristics from candidate segments and replacement target segments using the foundation model illustrated in Fig. 9 (Minz Won, et al., "A Foundation Model for Music Informatics", "IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024"). The foundation model is a massive basic model trained on an overwhelmingly large amount of data to perform various music tasks such as sound source separation, transcription, and genre classification, and further includes musical structural information.

[0045] The audio signal processing device can store the acquired musical characteristics as embedding information for each candidate segment and the segment to be replaced. In this case, the embedding information may be an embedding vector. Specifically, the audio signal processing device can generate an embedding vector of a segment based on the segment length, the segment energy distribution, or the segment musical characteristics.

[0046] FIG. 10 shows an audio signal processing device according to an embodiment of the present invention generating a replacement target segment from input music audio and generating embedding information including musical characteristics of the replacement target segment.

[0047] As previously explained, the audio signal processing device can acquire a replacement target segment containing music from the input music audio. At this time, the audio signal processing device can acquire the replacement target segment based on whether a signal containing music is detected from the input music audio or based on a change in the signal containing music. Additionally, the audio signal processing device can acquire only audio signals containing music as replacement target segments, excluding non-music sound effects and dialogue.

[0048] When an audio signal processing device acquires a segment to be replaced from an input music audio signal, if the difference between the end time of the first segment to be replaced and the start time of the second segment to be replaced is less than or equal to a reference value, the segments may be merged into a single segment to be replaced. In this case, the reference value may be a value specified in advance. In another specific embodiment, the reference value may be determined based on the length of the first segment to be replaced or the second segment to be replaced. Specifically, if the length of the first segment to be replaced or the second segment to be replaced increases, the magnitude of the reference value may increase. Through this, the audio signal processing device can prevent relatively long music from being judged as multiple segments.

[0049] This is because when fade-in or fade-out is used within a single song, it may be perceived as separate tracks due to changes in volume, even though it is actually one song. Additionally, if the length of the segment to be replaced acquired by the audio signal processing unit is greater than a preset value, the audio signal processing unit may advance the start point of the acquired segment to be replaced within the input music audio signal or delay the end point. This is to prevent the corresponding part from being excluded from the segment when fade-in or fade-out is used.

[0050] Additionally, the audio signal processing device can advance the start time of the acquired replacement segment or delay the end time of the input music audio signal based on the degree of reverb included in the replacement segment. For example, the reverberation time (RT60) of the replacement segment can be calculated, and the end can be delayed by a calculated extension time based on that time. Through this method, the audio signal processing device can adjust the length of the segment to be longer as the degree of reverb increases.

[0051] Additionally, the audio signal processing device can distinguish between different replacement target segments at a point where the change in musical characteristics differs more than a predetermined degree, even if it is a single piece of music. Additionally, the audio signal processing device can process multiple pieces of music as a single replacement target segment even if they are consecutive, provided that the change in musical characteristics is smaller than a predetermined degree. In the embodiments described above, the musical characteristics may include at least one of genre, instrument composition, tempo, or key.

[0052] The audio signal processing device can generate embedding information including musical characteristics of the segment to be replaced. The method by which the audio signal processing device generates embedding information including musical characteristics of the segment to be replaced may be the same as the method for generating embedding information of the candidate segment described above.

[0053] The audio signal processing device can determine a candidate segment to replace the segment to be replaced by comparing the embedding information of the segment to be replaced with the embedding information of candidate segments stored in the candidate segment database. For convenience of explanation, the candidate segment to replace the segment to be replaced is referred to as the final candidate segment. The audio signal device can determine the candidate segment having embedding information most similar to the embedding information of the segment to be replaced as the final candidate segment. At this time, the audio signal processing device can determine the similarity using at least one of the Euclidean distance or cosine similarity between the embedding vector of the segment to be replaced and the embedding vector of the candidate segment.

[0054] When an audio signal processing unit replaces a target segment with a final candidate segment, it is necessary to ensure that the overall progression of the input music audio signal does not sound awkward. To achieve this, a method is required for the audio signal processing unit to adjust the loudness of the final candidate segment.

[0055] In a specific embodiment, the audio signal processing device can determine the final candidate segment based on the embedding information of the segment to be replaced and the embedding information of the candidate segment, as well as the signal envelope contour of the segment to be replaced. For example, the audio signal processing device can determine the final candidate segment based on the similarity between the embedding information of the segment to be replaced and the embedding information of the candidate segment, and the similarity between the signal envelope contour of the segment to be replaced and the signal envelope contour of the candidate segment.

[0056] Additionally, when the audio signal processing device determines the final candidate segment, it can determine, based on user input, which of the factors used to judge similarity has the highest weight. In this case, the factors used to judge similarity may include any one of the similarity of the instrument configuration used, the similarity of the beat, the similarity of the melody, or the similarity of the mood of the song.

[0057] The audio signal processing device can estimate the mixing gain used when the music of the replacement target segment is inserted into the input music audio signal, and apply the estimated mixing gain to the final candidate segment. For example, the audio signal processing device can estimate the fade-in and fade-out curves applied to the replacement target segment, and apply the estimated fade-in and fade-out curves to the final candidate segment. The audio signal processing device can calculate loudness-related values ​​such as RMS, Momentary Loudness, and Short-term Loudness for the replacement target segment in short time units, distinguish the representative value of the accumulated values ​​as the gain and the characteristics of the distribution of the accumulated values ​​as the dynamic range characteristic, and apply the distinguished gain and dynamic range characteristic to the final candidate segment. Through this, the audio signal processing device can replace the replacement target segment with the final candidate segment without disrupting the loudness balance of the input music audio signal.

[0058] When an audio signal processing device replaces a segment to be replaced by considering only the average loudness level of the final candidate segment, it may obscure input non-musical audio signals, such as dialogue and sound effects, that need to be mixed after replacement. In particular, as the frequency bands of the signal of the segment to be replaced and the final candidate segment differ, even candidate segments with the same loudness level as the segment to be replaced may obscure input non-musical audio signals included in the audio signal.

[0059] The audio signal processing device can adjust the loudness of the final candidate segment by considering the spectrum or instantaneous loudness level of the input non-musical audio segment to be mixed. Specifically, the audio signal processing device can first adjust the average loudness level of the final candidate segment. The audio signal processing device can calculate a masking threshold value, which is the minimum value at which the input non-musical signal is masked by the adjusted loudness level, calculate a correction gain so that the masking threshold value does not obscure the entire input non-musical audio signal to be mixed or the dialogue signal among them, and then apply the correction gain to the final candidate segment to adjust the loudness level. At this time, the audio signal processing device can perform loudness normalization so that the instantaneous loudness level of the final candidate segment is smaller than the masking threshold value of the input non-musical audio signal.

[0060] Additionally, the audio signal processing device can acquire the instantaneous loudness of the segment to be replaced. At this time, the instantaneous loudness may include various methods such as short-term loudness or momentary loudness, or sample-by-sample gain based on the estimated energy ratio calculated per frame. The audio signal processing device can acquire a time-variable correction gain from the variable loudness and the variable loudness of the final candidate segment. The audio signal processing device can adjust the loudness of the final candidate segment by applying the acquired time-variable correction gain to the final candidate segment. According to a specific embodiment, the audio signal processing device may use the variable loudness of the speech signal of the segment to be replaced or the variable loudness of the output prediction signal, which is the result of mixing the input non-musical signal for the interval of the segment to be replaced and the segment to be replaced, instead of the variable loudness of the segment to be replaced.

[0061] When an audio signal processing device replaces a segment to be replaced with a final candidate segment, a difference in length between the segment to be replaced and the final candidate segment may be an issue. If the length of the final candidate segment is shorter than the length of the segment to be replaced, the audio signal processing device may append a portion of the final candidate segment to the front or back of the final candidate segment. In this case, the audio signal processing device may extend the length of the final candidate segment by a structural unit of music, such as a bar. In another specific embodiment, if the length of the final candidate segment is shorter than the length of the segment to be replaced, the audio signal processing device may extend the length of the final candidate segment by using pitch shifting or time stretching.

[0062] In addition, the audio signal processing device can divide a single song into multiple replacement target segments and independently determine a final candidate segment for each of the multiple replacement target segments. At this time, the length of the replacement target segment can be determined in units of the song's structure, such as in units of measures. Through this, the audio signal processing device can replace a single song with multiple songs and replace each section with the music having the highest similarity.

[0063] Some embodiments may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules executed by a computer. A computer-readable medium may be any available medium accessible by a computer and may include both volatile and non-volatile media, and both removable and non-removable media. Additionally, a computer-readable medium may include a computer storage medium. A computer storage medium may include both volatile and non-volatile, removable and non-removable media implemented by any method or technique for storing information, such as computer-readable instructions, data structures, program modules, or other data.

[0064] Although the present disclosure has been described above through specific embodiments, those skilled in the art with ordinary knowledge of the technical field to which the present disclosure pertains may make modifications and changes without departing from the spirit and scope of the present disclosure. That is, while the present disclosure describes embodiments of loudness level correction for audio signals, the present disclosure is equally applicable and extendable to various multimedia signals, including video signals as well as audio signals. Therefore, anything that can be easily inferred by a person skilled in the technical field to which the present disclosure pertains from the detailed description and embodiments of the present disclosure is interpreted as falling within the scope of the rights of the present disclosure.

Claims

Claim 1 An audio signal processing device for replacing music using a machine learning model, comprising a processor, wherein the processor receives an input music audio signal, obtains a replacement target segment including a signal containing music to be replaced from the input music audio signal, determines a final candidate target segment among the plurality of candidate segments based on the similarity between the plurality of candidate segments obtained by the machine learning model and the replacement target segment, and outputs an audio signal in which the replacement target segment in the input music audio signal is replaced with the final candidate target segment. Claim 2 An audio signal processing device according to claim 1, wherein the similarity between the plurality of candidate segments and the replacement target segment includes similarity of musical characteristics. Claim 3 An audio signal processing device according to claim 2, wherein the similarity between the plurality of candidate segments and the replacement target segment includes the similarity between the sound pressure envelope contour of the plurality of candidate segments and the sound pressure envelope contour of the replacement target segment. Claim 4 In claim 1, each of the plurality of candidate segments is a portion of music corresponding to each of the plurality of candidate segments, and the audio signal processing device stores information indicating a start time and an end time indicating the music corresponding to each of the plurality of candidate segments and the portion of music corresponding to each of the plurality of candidate segments. Claim 5 In claim 1, the processor is an audio signal processing device that adjusts the loudness of the final candidate target segment when replacing the replacement target segment with the final candidate target segment in the input music audio signal. Claim 6 In claim 5, the audio signal processing device, wherein the processor adjusts the loudness of the final candidate segment based on the spectrum or instantaneous loudness level of the replacement target segment when replacing the replacement target segment with the final candidate target segment in the input music audio signal. Claim 7 In claim 1, the audio signal processing device, when the processor replaces the replacement target segment in the input music audio signal with the final candidate target segment, if the length of the final candidate segment is shorter than the length of the replacement target segment, appends a part of the final candidate segment before or after the final candidate segment. Claim 8 An audio signal processing device according to claim 1, wherein the processor divides a song into a plurality of replacement target segments and independently determines the final candidate segment for each of the plurality of replacement target segments. Claim 9 In claim 1, the audio signal processing device, when the processor obtains the replacement target segment from the input audio signal, merges the first replacement target segment and the second replacement target segment into one replacement target segment if the difference between the end time of the first replacement target segment and the start time of the second replacement target segment is less than or equal to a reference value. Claim 10 A method of operation of an audio signal processing device for replacing music using a machine learning model, comprising: receiving an input music audio signal; obtaining a replacement target segment including a signal containing music to be replaced from the input music audio signal; determining a final candidate target segment among a plurality of candidate segments based on the similarity between a plurality of candidate segments obtained by the machine learning model and the replacement target segment; and outputting an audio signal in which the replacement target segment in the input music audio signal is replaced with the final candidate target segment. Claim 11 In claim 10, the similarity between the plurality of candidate segments and the replacement target segment includes similarity of musical characteristics. Claim 12 A method of operation according to claim 11, wherein the similarity between the plurality of candidate segments and the replacement target segment includes the similarity between the envelope contour of the sound pressure of the plurality of candidate segments and the envelope contour of the sound pressure of the replacement target segment. Claim 13 In claim 10, each of the plurality of candidate segments is a portion of music corresponding to each of the plurality of candidate segments, and the audio signal processing device stores information indicating a start time and an end time indicating the music corresponding to each of the plurality of candidate segments and the portion of music corresponding to each of the plurality of candidate segments. Claim 14 In claim 10, the above-described method of operation further comprises the step of adjusting the loudness of the final candidate target segment when replacing the replacement target segment with the final candidate target segment in the input music audio signal. Claim 15 In claim 14, the step of adjusting the loudness of the final candidate target segment further comprises the step of adjusting the loudness of the final candidate segment based on the spectrum or instantaneous loudness level of the replacement target segment when replacing the replacement target segment with the final candidate target segment in the input music audio signal. Claim 16 In claim 10, the operation method comprises the step of, when replacing the replacement target segment in the input music audio signal with the final candidate target segment, if the length of the final candidate segment is shorter than the length of the replacement target segment, appending a part of the final candidate segment before or after the final candidate segment. Claim 17 In claim 10, the step of determining the final candidate target segment among the plurality of candidate segments further comprises the step of dividing a single song into a plurality of replacement target segments and independently determining the final candidate segment for each of the plurality of replacement target segments. Claim 18 In claim 10, the step of acquiring the replacement target segment includes a step of merging into one replacement target segment when the difference between the end time of the first replacement target segment and the start time of the second replacement target segment is less than or equal to a reference value when acquiring the replacement target segment from the input audio signal.