AI pet implementation method based on multi-dimensional emotion model and personalized sound effect

By constructing AI pet personalities using four-dimensional dynamic parameters and the Plutchk emotion model, and combining this with acoustic parameter adjustments, personalized sound effects are generated, solving the problem of monotonous emotions in traditional electronic pets and improving the user experience.

CN121237104APending Publication Date: 2025-12-30呜呦(长沙)智能科技有限公司
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511130838.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Traditional electronic pets offer limited emotional expression and lack personalized and natural feedback; their voice systems fail to provide a truly personalized experience for each user.

Method used

The emotional state of AI pets is characterized by four-dimensional dynamic parameters. Personality characteristics are constructed based on the Plutchk emotion model. The original sound is adjusted by filter equations, frequency shifts, and formant shift parameters to generate personalized sound effects.

Benefits of technology

This makes the emotional expression of AI pets more realistic and personalized, improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237104A_ABST
    Figure CN121237104A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of AI pets, and discloses an AI pet implementation method based on a multi-dimensional emotion model and a personalized sound effect, and the method comprises the steps: representing the emotion state of an AI pet through employing a four-dimensional dynamic parameter; constructing character features of the AI pet through the four-dimensional dynamic parameters on the basis of a Puluxik emotion model; according to character characteristics, sequentially performing filter equation, frequency translation amount and formant offset parameter adjustment on the original sound subjected to short-time Fourier transform; and performing inverse short-time Fourier transform processing on the sound after parameter adjustment to generate the sound effect of the AI pet. Four-dimensional dynamic parameters are combined into various composite emotions through the Puluxik emotion model, the traditional binary emotion classification limitation is broken through, the character characteristics of the AI pet are constructed, and the original sound is subjected to parameter adjustment in combination with the character characteristics, so that the sound effect of the AI pet is generated. The fixed sound effect limitation is broken through, the personalized sound effect is realized, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of AI pet technology, specifically to a method for realizing AI pets based on multidimensional emotion models and personalized sound effects. Background Technology

[0002] With the development of technology, electronic pets have become widely used. These pets simulate the appearance and behavior of real pets through electronic devices, fulfilling users' needs for companionship and entertainment. However, traditional electronic pets typically interact with users based on preset behavioral patterns, resulting in limited emotional expression, a lack of personalized and natural feedback, and simplistic emotional models that rely on simple binary classifications such as happy or sad, failing to simulate complex emotional states. Furthermore, their sound systems often depend on fixed sound effect libraries, making it impossible to provide a personalized experience tailored to each user. Summary of the Invention

[0003] In view of this, the present invention provides an AI pet implementation method based on a multi-dimensional emotion model and personalized sound effects to solve the problems of traditional electronic pets having single emotions and fixed sound effects.

[0004] In a first aspect, the present invention provides a method for implementing an AI pet based on a multidimensional emotion model and personalized sound effects, the method comprising:

[0005] Four-dimensional dynamic parameters are used to characterize the emotional state of AI pets;

[0006] Based on the Plutschke emotion model, the personality traits of AI pets are constructed through four-dimensional dynamic parameters.

[0007] Based on personality traits, the original sound after short-time Fourier transform is adjusted sequentially by applying filter equations, frequency shifts, and formant shift parameters.

[0008] The sound after parameter adjustment is processed by inverse short-time Fourier transform to generate the sound effect of the AI ​​pet.

[0009] This invention combines four-dimensional dynamic parameters into various complex emotions using the Plutschke emotion model, breaking through the limitations of traditional binary emotion classification, constructing AI pet personality traits, and adjusting the parameters of the original voice based on these personality traits, thus overcoming the limitations of fixed sound effects, achieving personalized sound effects, and improving the user experience.

[0010] In one alternative implementation, the method further includes:

[0011] If no similar emotional stimulus is received during the emotional decay period, the intensity of the emotion should be gradually reduced.

[0012] If similar emotional stimuli are received during the period of emotional decline, the intensity of the emotion will be amplified;

[0013] If similar emotional stimuli occur consecutively within a preset time period, the current emotional tone will be strengthened.

[0014] If negative emotional stimuli occur within a preset time period, the emotional state will be changed.

[0015] This invention uses a mechanism of emotional decay and superposition of similar stimuli to adjust the intensity, tone, and state of emotions, simulating the emotional continuity of real organisms, thus making the emotional expression of AI pets more realistic.

[0016] In one alternative implementation, after constructing the personality traits of the AI ​​pet, the method further includes:

[0017] By using an exponentially weighted moving average model to repeatedly stimulate and learn from input emotional signals over a long period, the long-term personality of an AI pet can be constructed.

[0018] This invention utilizes an exponentially weighted moving average model to construct the long-term personality of an AI pet, allowing the AI ​​pet's personality to gradually stabilize during interaction with the user, forming a stable long-term personality tone.

[0019] In one optional implementation, adjustments are made to the filter equation, frequency shift, and formant shift parameters, including:

[0020] Adjust the frequency shift using the following formula:

[0021] X shifted (f)=X(ff shift )

[0022] Among them, f shift This represents the amount of frequency shift;

[0023] Adjust the filter equations according to the following formula:

[0024] X filtered (f)=H(f)X(f)

[0025] Where H is the input filter equation;

[0026] Adjust the resonance peak shift according to the following formula:

[0027] X warped (f) = X(f·α)

[0028] Where α is a parameter representing the resonant peak shift.

[0029] This invention achieves adaptive adjustment of acoustic parameters by adjusting the frequency offset to change the pitch of the sound, adjusting the filter equation to change the frequency content of the sound, and adjusting the formant offset to change the timbre of the sound.

[0030] In one alternative implementation, the personality traits include cheerful, active, shy, and clingy, and the filter equation is adjusted as follows:

[0031] When the personality type is outgoing, a high-pass filter is selected, and the Fourier domain equation is as follows:

[0032]

[0033] Among them, f c =400Hz, n varies according to the value of the light;

[0034] When the personality type is active, a bandpass filter is selected, and the Fourier domain equation is as follows:

[0035]

[0036] in, n changes according to the active value;

[0037] When the personality type is shy, a low-pass filter is selected, and the Fourier domain equation is as follows:

[0038]

[0039] Among them, f c =280Hz, n and G change according to the value of shy;

[0040] When the personality type is clingy, a bandpass filter is selected, and the Fourier domain equation is as follows:

[0041]

[0042] in, n and G change according to the value of stickiness.

[0043] This invention selects corresponding filters for different personality types and sets parameters to make the sound feedback match the corresponding personality type, allowing users to intuitively distinguish the personality of the AI ​​pet through sound.

[0044] In one alternative implementation, the method further includes, before generating the AI ​​pet's sound effects:

[0045] The sound band range is preset, which is divided into low band, medium band and high band;

[0046] Select the target band within the pre-defined sound band range that corresponds to the AI ​​pet species type.

[0047] This invention filters out harsh sounds by pre-setting the sound frequency range and selects the frequency range that matches the species of the AI ​​pet within the preset sound frequency range, making the AI ​​pet's voice more similar to that of a real pet and improving the user experience.

[0048] In one alternative implementation, generating the AI ​​pet's sound effects includes:

[0049] Pre-set basic sound feedback;

[0050] Based on basic sound feedback and combined with personality traits, the AI ​​pet's sound effects are dynamically generated.

[0051] This invention provides a unified basic sound effect by setting a basic sound effect feedback, and on this basis, combined with personality traits, generates unique sound effects for AI pets, thus achieving personalized sound feedback.

[0052] Secondly, this invention provides an AI pet implementation system based on a multi-dimensional emotion model and personalized sound effects, the system comprising:

[0053] The emotion representation module is used to represent the emotional state of the AI ​​pet using four-dimensional dynamic parameters.

[0054] The personality building module is used to construct the personality traits of AI pets based on the Plutschke emotion model through four-dimensional dynamic parameters.

[0055] The parameter adjustment module is used to adjust the filter equation, frequency shift, and formant shift parameters of the original sound after short-time Fourier transform according to personality characteristics.

[0056] The sound effects generation module is used to perform inverse short-time Fourier transform processing on the adjusted sound to generate sound effects for the AI ​​pet.

[0057] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the AI ​​pet implementation method based on a multidimensional emotion model and personalized sound effects described in the first aspect or any corresponding embodiment above.

[0058] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the AI ​​pet implementation method based on a multidimensional emotion model and personalized sound effects described in the first aspect or any corresponding embodiment above. Attached Figure Description

[0059] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0060] Figure 1 This is a flowchart illustrating the method for implementing an AI pet based on a multidimensional emotion model and personalized sound effects according to an embodiment of the present invention.

[0061] Figure 2 This is a schematic diagram of 32 emotions based on 8 basic components according to an embodiment of the present invention;

[0062] Figure 3 This is a personality diagram according to an embodiment of the present invention;

[0063] Figure 4 This is a schematic diagram of the sound processing flow according to an embodiment of the present invention;

[0064] Figure 5 This is a structural block diagram of an AI pet implementation system based on a multidimensional emotion model and personalized sound effects according to an embodiment of the present invention;

[0065] Figure 6 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0067] According to an embodiment of the present invention, an embodiment of an AI pet implementation method based on a multidimensional emotion model and personalized sound effects is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0068] This embodiment provides a method for implementing an AI pet based on a multi-dimensional emotion model and personalized sound effects, which can be widely applied to fields such as smart toys and companion robots. Figure 1This is a flowchart of an AI pet implementation method based on a multi-dimensional emotion model and personalized sound effects according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps:

[0069] Step S101: Use four-dimensional dynamic parameters to characterize the emotional state of the AI ​​pet.

[0070] In this embodiment of the invention, the basic emotional state of the AI ​​pet is expressed through four-dimensional dynamic parameters. These four-dimensional dynamic parameters are paired and opposite to each other; four numbers from -1 to 1 represent the parameters, such as... Figure 2 As shown, the eight basic emotions are sadness, happiness, disgust, trust, surprise, anticipation, fear, and anger, which can form 32 different complex emotions.

[0071] Specifically, the four-dimensional dynamic parameters are shown in Table 1:

[0072] Table 1

[0073] Emotion Group -1 1 No. 1 sad hapiness No. 2 disgust trust No. 3 surprise expect No. 4 fear angry

[0074] Step S102: Based on the Plutschke emotion model, the four-dimensional dynamic parameters are combined to construct the personality traits of the AI ​​pet.

[0075] In this embodiment of the invention, based on Plutchik's Emotion Model, complex personality traits of the AI ​​pet are constructed through four-dimensional dynamic parameters. These personality traits include cheerfulness, activity, shyness, and clinginess.

[0076] Specifically, preset personality types are generated through the combination of four-dimensional dynamic parameters. The personality definition rules are shown in Table 2 below, which are only examples and not intended to be limiting:

[0077] Table 2

[0078]

[0079] like Figure 3 As shown, Figure 3 This is a diagram illustrating a personality type.

[0080] Step S103: Based on personality traits, the original sound after short-time Fourier transform is adjusted sequentially by applying filter equations, frequency shifts, and formant shift parameters.

[0081] In this embodiment of the invention, an adaptive acoustic parameter adjustment algorithm is introduced, incorporating personality traits, such as... Figure 4As shown, the original sound is first subjected to a short-time Fourier transform, and then the filter equation, frequency shift, and formant shift parameters are adjusted in sequence. Adjusting the filter equation changes the frequency content of the sound, adjusting the frequency shift changes the pitch of the sound, and adjusting the formant shift makes the sound sound like another person.

[0082] Step S104: Perform inverse short-time Fourier transform on the adjusted sound to generate the sound effect of the AI ​​pet.

[0083] In this embodiment of the invention, after the sound is adjusted by parameters, it undergoes inverse short-time Fourier transform processing to obtain a new sound, generating the sound effects of the AI ​​pet and giving each AI pet a unique sound.

[0084] The AI ​​pet implementation method based on multidimensional emotion model and personalized sound effects provided in this embodiment combines four-dimensional dynamic parameters into various complex emotions through the Plutschke emotion model, breaking through the limitations of traditional binary emotion classification, constructing AI pet personality characteristics, and combining personality characteristics to adjust the parameters of the original sound, breaking through the limitations of fixed sound effects, realizing personalized sound effects, and improving the user experience.

[0085] This embodiment provides a method for implementing an AI pet based on a multi-dimensional emotion model and personalized sound effects. The process includes the following steps:

[0086] Step S201: Use four-dimensional dynamic parameters to characterize the emotional state of the AI ​​pet.

[0087] Please see details Figure 1 Step S101 of the illustrated embodiment will not be described again here.

[0088] In some alternative implementations, the method further includes:

[0089] Step S202: If no similar emotional stimulus is received during the emotional decay period, the emotional intensity is gradually reduced.

[0090] Step S203: If the same emotional stimulus is received during the emotional decay period, the emotional intensity is increased.

[0091] Step S204: If similar emotional stimuli occur continuously within a preset time period, the current emotional tone is strengthened.

[0092] Step S205: If negative emotional stimuli occur within a preset time period, change the emotional state.

[0093] In this embodiment of the invention, an emotion dynamic regulation mechanism is introduced, including an emotion decay mechanism and an emotion memory and superposition mechanism.

[0094] Specifically, emotional stimuli such as touch and voice interaction from users will cause the AI ​​pet to generate emotions of a certain intensity. However, the intensity of this emotion is not fixed. If no similar emotional stimulus is received during the emotion decay period, the intensity of the emotion will gradually decrease. When a similar emotional stimulus is received during the emotion decay period, the intensity of the emotion will increase and be superimposed to achieve a smooth transition of emotions.

[0095] The system records user interactions with the AI ​​pet within a preset timeframe. Taking ten minutes as an example, if similar emotional stimuli occur consecutively within ten minutes, the current emotional tone is reinforced. For instance, if the user interacts repeatedly with a cheerful tone, the "happy" emotion is triggered multiple times. If negative emotional stimuli occur within the preset timeframe, an emotional state transition is triggered. For example, a loud scolding will transition to a "sad or fear" emotion.

[0096] By employing mechanisms of emotional decay and superposition of similar stimuli, the intensity, tone, and state of emotions are adjusted to simulate the emotional continuity of real organisms, making the emotional expression of AI pets more realistic.

[0097] Step S206: Based on the Plutschke emotion model, construct the personality traits of the AI ​​pet through four-dimensional dynamic parameters.

[0098] Please see details Figure 1 Step S102 of the illustrated embodiment will not be described again here.

[0099] In some alternative implementations, the method further includes:

[0100] Step S207: Use an exponentially weighted moving average model to repeatedly stimulate and learn the input emotional signals over a long period of time to construct the long-term personality of the AI ​​pet.

[0101] In this embodiment of the invention, during the interaction between the user and the AI ​​pet, the AI ​​pet will be influenced by the user, thereby constructing a long-term personality.

[0102] Specifically, the AI ​​pet's long-term personality is constructed by using an Exponential Weighted Moving Average (EWMA) model to repeatedly stimulate and learn from the input emotional signals.

[0103] The exponentially weighted moving average model is as follows:

[0104]

[0105] Among them, EWMA t The EWMA value at sampling time point t. t-1 The EWMA value is the value at the previous sampling point t-1. For the learning rate, r tInput the most recently received emotional signal.

[0106] An exponentially weighted moving average model is used to construct the long-term personality of an AI pet, so that the AI ​​pet's personality gradually stabilizes during the interaction with the user, forming a stable long-term personality tone.

[0107] Step S208: Based on personality traits, the original sound after short-time Fourier transform is adjusted sequentially by applying filter equations, frequency shifts, and formant shift parameters.

[0108] Specifically, the adjustments to the filter equation, frequency shift, and resonant shift parameters in step S208 above include:

[0109] Step S2081: Adjust the frequency shift according to the following formula:

[0110] X shifted (f)=X(ff shift )

[0111] Among them, X shifted The sound after being processed by the frequency converter, f shift This represents the amount of frequency shift.

[0112] Step S2082: Adjust the filter equation according to the following formula:

[0113] X filtered (f)=H(f)X(f)

[0114] Among them, X filtered H represents the sound after filtering, and H is the input filter equation.

[0115] Step S2083: Adjust the resonance peak shift according to the following formula:

[0116] X warped (f) = X(f·α)

[0117] Among them, X warped The sound is after resonance shift processing, where α is the parameter for resonant peak shift.

[0118] In this embodiment of the invention, adjusting the filter equation H can alter the frequency content of the sound. Specifically, removing high frequencies can make the sound sound deeper or more distant. Removing low frequencies can make the sound sound thinner or sharper. Emphasizing a specific frequency range can make the sound clearer and sharper.

[0119] The amount of frequency shift f shift It can change the pitch of a sound, making it sharper or lower; the amount of frequency shift f is determined by the activity value. shift .

[0120] The formant shift parameter α can make a voice sound like a different person, and the formant shift parameter α is determined according to the degree of stickiness. Specifically, increasing the resonant frequency will make the speaker's voice sound softer (like a child or a chipmunk), while decreasing the resonant frequency will make the speaker's voice sound louder (like a giant).

[0121] Steps S2081 to S2083 process the same audio file. These three steps can be combined into one step, as shown in the following formula:

[0122] X filtered (f)=H(f·α-f shift )X(f·α-f shift )

[0123] By adjusting the frequency offset, the pitch of the sound is changed; by adjusting the filter equation, the frequency content of the sound is changed; and by adjusting the formant offset, the timbre of the sound is changed, thus achieving adaptive adjustment of acoustic parameters.

[0124] Specifically, the personality traits include cheerful, active, shy, and clingy. Step S2082 above, adjusting the filter equation, includes:

[0125] Step S2082a: When the personality type is cheerful, a high-pass filter is selected, and the Fourier domain equation is as follows:

[0126]

[0127] Among them, f c =400Hz, n varies according to the value of the light.

[0128] Step S2082b: When the personality type is active, a bandpass filter is selected, and the Fourier domain equation is as follows:

[0129]

[0130] in, n changes according to the active value.

[0131] Step S2082c: When the personality type is shy, a low-pass filter is selected, and the Fourier domain equation is as follows:

[0132]

[0133] Among them, f c =280Hz, n and G change according to the value of shy.

[0134] Step S2082d: When the personality type is clingy, a bandpass filter is selected, and the Fourier domain equation is as follows:

[0135]

[0136] in, n and G change according to the value of stickiness.

[0137] In this embodiment of the invention, personality types include cheerful (E1), active (E2), shy (E3), and clingy (E4).

[0138] Specifically, when the personality type is cheerful, increasing the formant shift frequency can make the sound younger and more exciting, increasing the pitch upward can make the sound more cheerful, and boosting the high frequencies (high-pass filtering) can increase energy and brightness.

[0139] When the personality type is active, bandpass filtering can highlight the mid-frequency range, making the voice / instrument clearer and more distinct, while a slight frequency shift will make the sound more rapid.

[0140] When the personality type is shy, a low-pass filter can remove sharp high frequencies, making the sound deeper and more resonant. Lower resonance makes the sound softer and more direct, and a slight frequency shift makes the sound feel smaller.

[0141] When the personality type is clingy, enhancing the mid-range frequencies will make the voice sound more intimate. Increasing the resonant frequency will produce a nasal or high-resonance effect, making the user feel intimate. No frequency shift (keeping the pitch stable) ensures that the voice feels familiar and intimate to the user.

[0142] By selecting corresponding filters and setting parameters for different personality types, the sound feedback can be made to match the corresponding personality type, allowing users to intuitively distinguish the personality of the AI ​​pet through its voice.

[0143] In some alternative implementations, the method further includes:

[0144] Step S209: Pre-set the sound band range, which is divided into low band, medium band and high band.

[0145] Step S210: Select the target band from the preset sound band range that corresponds to the AI ​​pet species type.

[0146] In this embodiment of the invention, the hearing range of an average person is 20-20000Hz. In order to avoid harsh sounds, the sound band range is preset, and the sound of the 50-1500Hz band is selected and divided into 3 intervals: low band: 50-280Hz, medium band: 281-399Hz, and high band: 400Hz.

[0147] Different species can produce different ranges of sounds. The target band is selected from a pre-defined range of sound frequencies that match the species type of the AI ​​pet. Specifically, humans can produce sounds in the range of 85-1100Hz, cats in the range of 760-1520Hz, and dogs in the range of 450-1080Hz. For example, if the AI ​​pet's species type is a dog, then the sound frequency range of dogs will be selected.

[0148] By pre-setting the sound frequency range and filtering out harsh sounds, the AI ​​pet's voice is made to sound more like a real pet's voice, thus improving the user experience.

[0149] Step S211: Perform inverse short-time Fourier transform on the adjusted sound to generate the sound effect of the AI ​​pet.

[0150] Specifically, the sound effects generated for the AI ​​pet in step S211 above include:

[0151] Step S211a: Preset basic sound effect feedback.

[0152] Step S211b: Based on basic sound effect feedback and combined with personality traits, dynamically generate the sound effects of the AI ​​pet.

[0153] In this embodiment of the invention, five basic sound effects are preset: singing, happy, calm, tense, and low battery alarm. After the basic sound effects are set, natural expression can be achieved by randomly adjusting the pitch, rhythm, and volume parameters.

[0154] Based on basic sound feedback and combined with personality traits, AI pet sound effects are dynamically generated. For example, cheerfulness corresponds to high-frequency tones.

[0155] Specifically, the AI ​​pet's sound effect triggering rules are as follows: when the emotional intensity reaches a threshold, the corresponding sound effect is triggered. For example, high-intensity "anger" triggers a high-frequency, sharp sound effect. After the cumulative number of emotional triggers reaches a preset number, basic sound effect feedback is activated. For example, after the cumulative triggering of the "happy" emotion 10 times, the "singing" feedback is activated.

[0156] The AI ​​pet implementation method based on a multi-dimensional emotion model and personalized sound effects provided in this embodiment provides a unified basic sound effect by setting basic sound effect feedback. On this basis, combined with personality traits, unique sound effects are generated for the AI ​​pet to achieve personalized sound feedback.

[0157] As a specific application embodiment of this invention, the AI ​​pet implementation method based on a multi-dimensional emotion model and personalized sound effects can be implemented in hardware. Specifically, the AI ​​pet incorporates a multi-modal sensor, a central processing unit, an audio device, and a vibration motor. The multi-modal sensor collects information including tactile, auditory, and kinematic data, and collects environmental input in real time, including touch pressure and voice commands. The emotion engine in the central processing unit calculates the current emotion intensity based on the environmental input, updates dynamic thought parameters, and generates a complex emotion. The audio device includes a speaker and a voice changer, which calls a preset sound source or the user's sound source, and dynamically synthesizes sound effects by combining personality traits. The vibration motor triggers corresponding physical feedback based on the emotional state, achieving a closed-loop interaction between emotion perception and physical feedback.

[0158] This embodiment also provides an AI pet implementation system based on a multi-dimensional emotion model and personalized sound effects. This system is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the system described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0159] This embodiment provides an AI pet implementation system based on a multi-dimensional emotion model and personalized sound effects, such as... Figure 5 As shown, it includes:

[0160] The emotion representation module 501 is used to represent the emotional state of the AI ​​pet using four-dimensional dynamic parameters.

[0161] The personality construction module 502 is used to construct the personality characteristics of AI pets based on the Plutschke emotion model through four-dimensional dynamic parameters.

[0162] The parameter adjustment module 503 is used to adjust the filter equation, frequency shift, and formant shift parameters of the original sound after short-time Fourier transform according to personality characteristics.

[0163] The sound effect generation module 504 is used to perform inverse short-time Fourier transform processing on the sound after parameter adjustment to generate the sound effect of the AI ​​pet.

[0164] In some alternative implementations, the system further includes:

[0165] The emotion intensity reduction module is used to gradually reduce the emotion intensity if no similar emotional stimulus is received during the emotion decay period.

[0166] The emotion intensity enhancement module is used to enhance the emotion intensity if the same emotional stimulus is received during the emotion decay period.

[0167] The mood enhancement module is used to strengthen the current mood if similar emotional stimuli occur continuously within a preset time.

[0168] The emotional state change module is used to change the emotional state if negative emotional stimuli occur within a preset time.

[0169] In some alternative implementations, the system further includes:

[0170] The long-term personality construction module is used to learn the long-term personality of AI pets by repeatedly stimulating the input emotional signals using an exponentially weighted moving average model.

[0171] In some alternative implementations, the parameter adjustment module 503 includes:

[0172] The first adjustment unit is used to adjust the frequency shift according to the following formula:

[0173] X shifted (f)=X(ff shift )

[0174] Among them, X shifted The sound after being processed by the frequency converter, f shift This represents the amount of frequency shift.

[0175] The second adjustment unit is used to adjust the filter equations according to the following formula:

[0176] X filtered (f)=H(f)X(f)

[0177] Among them, X filtered H represents the sound after filtering, and H is the input filter equation.

[0178] The third adjustment unit is used to adjust the resonance peak shift according to the following formula:

[0179] X warped (f) = X(f·α)

[0180] Among them, X warped The sound is after resonance shift processing, where α is the parameter for resonant peak shift.

[0181] In some alternative implementations, the second adjustment unit includes:

[0182] The first adjustment subunit is used to select a high-pass filter when the personality type is cheerful. The Fourier domain equation is as follows:

[0183]

[0184] Among them, fc =400Hz, n varies according to the value of the light.

[0185] The second adjustment subunit is used to select a bandpass filter when the personality type is active. The Fourier domain equation is as follows:

[0186]

[0187] in, n changes according to the active value.

[0188] The third adjustment subunit is used to select a low-pass filter when the personality type is shy. The Fourier domain equation is as follows:

[0189]

[0190] Among them, f c =280Hz, n and G change according to the value of shy.

[0191] The fourth adjustment subunit is used to select a bandpass filter when the personality type is clingy. The Fourier domain equation is as follows:

[0192]

[0193] in, n and G change according to the value of stickiness.

[0194] In some alternative implementations, the system further includes:

[0195] The sound band range setting module is used to preset the sound band range, which is divided into low band, medium band and high band.

[0196] The band selection module is used to select the target band from the preset sound band range that corresponds to the AI ​​pet species type.

[0197] In some alternative implementations, the sound effects generation module 504 includes:

[0198] The sound feedback setting unit is used to preset the basic sound feedback.

[0199] The sound effect generation unit is used to dynamically generate the sound effects of AI pets based on basic sound effect feedback and personality traits.

[0200] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0201] In this embodiment, the AI ​​pet implementation system based on multidimensional emotion models and personalized sound effects is presented in the form of functional units. Here, a unit refers to an ASIC (Application Specific Integrated Circuit), a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0202] This invention also provides a computer device having the above-described features. Figure 5 The system shown is an AI pet implementation based on a multi-dimensional emotion model and personalized sound effects.

[0203] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 6 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 6 Take a processor 10 as an example.

[0204] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0205] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0206] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0207] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0208] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.

[0209] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touch screen. Output device 40 may include a display device, etc.

[0210] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0211] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0212] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope of this application.

Claims

1. An AI pet implementation method based on a multi-dimensional emotion model and personalized sound effects, characterized in that, The method comprises: adopting four-dimensional dynamic parameters to represent the emotional state of the AI pet; constructing the personality characteristics of the AI pet based on the four-dimensional dynamic parameters and the Plutchik emotional model; according to the personality characteristics, sequentially performing filter equation, frequency shift amount, and formant shift parameter adjustment on the original sound after short-time Fourier transform; performing inverse short-time Fourier transform processing on the sound after parameter adjustment to generate the sound effect of the AI pet.

2. The method of claim 1, wherein, The method further comprises: if no similar emotional stimulus is received during the emotional decay period, gradually reducing the emotional intensity; if a similar emotional stimulus is received during the emotional decay period, enhancing the emotional intensity; if a similar emotional stimulus continuously appears within a preset time length, strengthening the current emotional keynote; if a negative emotional stimulus appears within a preset time length, changing the emotional state.

3. The method of claim 1, wherein, After constructing the personality characteristics of the AI pet, the method further comprises: using an exponentially weighted moving average model to perform long-term repeated stimulus learning on the input emotional signal to construct the long-term personality of the AI pet.

4. The method of claim 1, wherein, The parameter adjustment of the filter equation, the frequency shift amount, and the formant shift comprises: adjusting the frequency shift amount according to the following formula: X shifted (f) = X(f - f shift ) where X shifted is the sound after processing by the frequency transformer, f shift is the amount of frequency translation; adjusting the filter equation according to the following formula: X filtered (f) = H(f) X(f) where X filtered is the sound after filter processing, and H is the input filter equation. adjusting the formant shift according to the following formula: X warped (f) = X(f · a) where X warped is the sound after the resonance shift processing, and a is a parameter of the resonance peak shift.

5. The method of claim 4, wherein, The types of the personality characteristics include cheerful, active, shy, and clingy, and the adjustment of the filter equation comprises: when the personality type is cheerful, a high-pass filter is selected, and the Fourier domain equation is as follows: where f c = 400 Hz, n varies according to the cheerful value; when the personality type is active, a band-pass filter is selected, and the Fourier domain equation is as follows: wherein, n is transformed according to the value of active; when the personality type is shy, a low-pass filter is selected, and the Fourier domain equation is as follows: where f c = 280 Hz, n and G are varied according to the value of shyness; when the personality type is clingy, a band-pass filter is selected, and the Fourier domain equation is as follows: wherein n and G are transformed according to the values of the stickiness.

6. The method of claim 1, wherein, Before generating the sound effect of the AI pet, the method further comprises: pre-setting a sound waveband range, which is divided into a low waveband, a middle waveband, and a high waveband; selecting a waveband corresponding to the species type of the AI pet as a target waveband within the pre-set sound waveband range.

7. The method of claim 1, wherein, The generation of the sound effect of the AI pet comprises: pre-setting a basic sound effect feedback; based on the basic sound effect feedback, dynamically generating the sound effect of the AI pet in combination with the personality characteristics. 8.An AI pet implementation system based on a multi-dimensional emotion model and a personalized sound effect, characterized in that, The system comprises: an emotional representation module for adopting four-dimensional dynamic parameters to represent the emotional state of the AI pet; a personality construction module for constructing the personality characteristics of the AI pet based on the four-dimensional dynamic parameters and the Plutchik emotional model; a parameter adjustment module for sequentially performing filter equation, frequency shift amount, and formant shift parameter adjustment on the original sound after short-time Fourier transform according to the personality characteristics; a sound effect generation module for performing inverse short-time Fourier transform processing on the sound after parameter adjustment to generate the sound effect of the AI pet.

9. A computer device, comprising: It comprises: a memory and a processor, which are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the AI pet implementation method based on the multi-dimensional emotional model and the individualized sound effect according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing a computer to execute the AI pet implementation method based on the multi-dimensional emotion model and the personalized sound effect according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for converting emotional speech by combining rhythm parameters with tone parameters

    CN102184731A

  • Emotional voice data conversion method and device, computer equipment and storage medium

    CN112466314A

  • Robot emotion expression system and method

    CN114003643A

  • Intelligent method and system for virtual object social interaction and emotion feedback

    CN118860162A

  • Speech synthesis method and device based on hierarchical emotion distribution, equipment and medium

    CN119207372A