Method and apparatus for customized character voice conversion and video generation using artificial intelligence

The method addresses the limitations of conventional voice conversion by deriving voice parameters from character attributes to create natural, emotionally expressive videos, applicable to a wide range of characters, including those without vocal organs, enhancing user experience and content diversity.

KR1020260112939APending Publication Date: 2026-07-21INTELLECTURE FUTURE IP MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
INTELLECTURE FUTURE IP MANAGEMENT CO LTD
Filing Date
2026-07-01
Publication Date
2026-07-21

Smart Images

  • Figure PAT00001_ABST
    Figure PAT00001_ABST
Patent Text Reader

Abstract

The present invention relates to a method and apparatus for customized character voice conversion and video generation using artificial intelligence. The method comprises: (a) acquiring a visual model of a target character, attribute information of the character, and original voice data of a user; (b) determining voice conversion parameters by analyzing the attribute information of the character; (c) converting the original voice data of the user into target voice data corresponding to the characteristics of the character using the determined voice conversion parameters and an artificial intelligence-based voice conversion model; and (d) generating a video by combining the acquired visual model of the character and the target voice data. According to the present invention, the spontaneity, emotion, and individuality of the user's actual speech are maintained while the voice is automatically modulated to match the characteristics of any character. Furthermore, acoustic identity can be assigned based solely on the visual and physical attributes of non-existent virtual characters, robots, objects, fantasy creatures, etc., for which actual voice samples of the target speaker are absent.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a method and apparatus for customized character voice conversion and video generation using artificial intelligence, and more specifically, to a technology that provides a user with the experience of auditorily possessing and speaking of an arbitrary character by obtaining a visual model and attribute information of a target character, converting original voice data spoken by a user into target voice data corresponding to the characteristics of the character according to voice conversion parameters derived from the attribute information of the character, and then generating a video by combining the visual model of the character and the target voice data. Background Technology

[0002] With the recent development of artificial intelligence technology, technology that converts a user's voice or text into a character's voice or applies lip-sync to a character image to create talking character videos has become widely known in the field of character-based content creation.

[0003] As a category of conventional technology, Text-to-Speech (TTS)-based multi-speaker speech synthesis technology is known. This method recognizes the attributes of a character (gender, age, etc.) from text information and selects a pre-stored speaker corresponding to them to read the text. However, this method has a fundamental limitation in that it leaves no room for the user's actual speech to intervene and merely maps a pre-prepared text to one of a finite number of pre-stored speakers to read it, thus failing to reflect the spontaneity, emotion, and individuality of the user's real-time speech.

[0004] Another category of conventional technology is Voice Conversion (VC) technology, which converts the voice of a source speaker into the acoustic characteristics of a target speaker. This method involves transplanting only the acoustic characteristics of the target speaker while maintaining the content of the source speaker's speech. However, conventional voice conversion technology has the limitation that learning and conversion are only possible if a large number of actual voice samples of the target speaker are secured in advance. Consequently, there is a problem in that it is difficult to apply to non-existent virtual characters, newly created characters, non-biological objects that do not possess actual vocal organs such as robots, dolls, or objects, or fantasy creatures, as it is impossible to secure the target voice samples themselves.

[0005] In addition, conventional voice-driven character animation technology receives a fixed audio signal as input and applies lip-syncing to the character image, but does not include the process of modulating the audio signal to match the characteristics of the character. Consequently, when the user's original voice is used as is, there is a problem in that the visual appearance of the character (e.g., a giant robot, a small fairy, a metallic object) and the auditory characteristics (the user's natural voice as an adult male or female) are severely discordant, causing cognitive dissonance in the user.

[0006] Furthermore, conventional lip-sync technology is limited to the lip movements of humanoid characters, so it is impossible to express visual speech for robots, dolls, holograms, or abstract characters that do not have lips, or it has limitations in that the speech expression does not match the physical characteristics of the character because only uniform lip shape animation is added. The problem to be solved

[0007] The present invention has been devised to solve the problems of the prior art described above, and aims to provide an AI-based customized character voice conversion and video generation method and apparatus that automatically modulates the voice to match the characteristics of an arbitrary character while maintaining the spontaneity, emotion, and individuality of the user's actual speech, and further combines the modulated voice with a visual model of the character to output it in the form of a natural video.

[0008] In addition, the present invention has another objective of enabling the derivation of voice conversion parameters using only the visual and physical attribute information of a character even in the absence of an actual voice sample of a target speaker, thereby enabling natural character voices to be assigned to non-existent virtual characters, robots, dolls, objects, fantasy creatures, etc.

[0009] In addition, another objective of the present invention is to fundamentally resolve the cognitive dissonance between visual appearance and auditory characteristics by automatically mapping physical attributes such as the external size, constituent materials, and internal cavity structure of a character to acoustic parameters such as vocal tract length, impulse response, and Helmholtz resonance.

[0010] In addition, another objective of the present invention is to provide a lip-sync method that utilizes alternative visual elements, such as a speaker grille, a light-emitting part, a vibration part, and a waveform display part, as the speech part, so that speech expression is possible even for characters that do not have lips.

[0011] In addition, another objective of the present invention is to extract the emotional state contained in the user's speech and reflect it in the character's voice, thereby enabling the user's original emotions to be naturally conveyed even after being converted into a character's voice. means of solving the problem

[0012] To solve the above-mentioned problem, a method for customized character voice conversion and video generation using artificial intelligence according to one aspect of the present invention comprises: (a) acquiring a visual model (10) of a target character, attribute information (20) of the character, and original voice data (30) of a user, performed by a computing device; (b) determining voice conversion parameters (40) by analyzing the attribute information (20) of the character; (c) converting the original voice data (30) of the user into target voice data (50) corresponding to the characteristics of the character using the determined voice conversion parameters (40) and an artificial intelligence-based voice conversion model (400); and (d) generating a video (60) by combining the acquired visual model (10) of the character and the target voice data (50).

[0013] A customized character voice conversion and video generation device (100) using artificial intelligence according to another aspect of the present invention comprises: a memory (120) for storing commands; and a processor (110) for executing commands, wherein the processor (110) acquires a visual model (10) of the character, attribute information (20), and original voice data (30) of the user, and converts the original voice data (30) into target voice data (50) based on voice conversion parameters (40) determined by analyzing the attribute information (20), and then combines the visual model (10) and the target voice data (50) to generate a video (60). Effects of the invention

[0014] According to the present invention, while the spontaneity, emotion, and individuality contained in the user's actual speech are preserved, only the acoustic characteristics of the speech are modulated and output to match an arbitrary character, thereby providing a new user experience in which the user can auditorily possess an arbitrary character and speak. This is a fundamentally different technical approach from the method of mapping a pre-prepared text to one of a finite number of pre-registered speakers to read aloud, and it exhibits a synergistic effect that cannot be achieved with conventional technology in that the user's real-time speech is freely expressed through a medium called a character.

[0015] In addition, according to the present invention, voice conversion parameters are automatically derived solely from the visual and physical attribute information of a character, without relying on the actual voice sample of a target speaker. Conventional voice conversion technology has an inherent limitation in that learning and conversion are possible only if the actual voice sample of a target speaker is secured in advance, making it fundamentally impossible to apply to non-existent characters, newly created characters, non-biological objects without vocal organs, or fantasy creatures. The present invention achieves a technological advancement that expands the scope of application of voice conversion technology itself by overcoming this fundamental limitation and enabling acoustic identity to be assigned solely from the visual and physical attributes of any object for which an actual voice sample does not exist in the world.

[0016] Furthermore, according to the present invention, voice conversion parameters are determined based on acoustically established physical principles, such as the physical law that the external size of a character is inversely proportional to the length of the vocal tract, inherent resonance characteristics based on the material, and the Helmholtz resonance principle based on internal cavities. Accordingly, cognitive dissonance between visual appearance and auditory characteristics is fundamentally resolved, and users can experience immersive content in which the character's visual appearance and the spoken voice naturally align. Compared to conventional methods that involve individually recording voices for each character or manually constructing pre-mapping tables, this invention possesses significant industrial applicability in that a voice conforming to physical laws is automatically assigned based solely on the physical attributes of any new character, even if the character is created on the spot.

[0017] In addition, according to the present invention, natural lip-syncing is possible for any character, such as a robot, doll, object, hologram, or abstract shape that does not have lips, by utilizing alternative visual elements such as a speaker grille, a light-emitting part, a vibration part, a waveform display part, or a shape-changing part as the speech part. This has the effect of dramatically expanding the range of expression in the character content industry in that it encompasses a vast area of ​​non-human characters that conventional human-face lip-sync technology could not cover in principle.

[0018] In addition, according to the present invention, an emotional state is extracted from the user's original voice data and reflected in the intonation and timbre of the target voice data, so the user's original emotions, such as joy, sadness, anger, and surprise, are naturally conveyed even after being converted into a character's voice. Accordingly, the character voice does not degenerate into a mechanical recitation with emotional voids, and a new channel for emotional expression is opened in which the user's emotions are conveyed to the listener in a reinterpreted form after passing through the filter of the character.

[0019] In addition, according to the present invention, the audiovisual naturalness of the final output is automatically guaranteed within the system through a feedback structure that calculates a cross-modal alignment score between a visual model of a character and target voice data and readjusts parameters when the alignment is below a threshold. This enables the system to verify and correct audiovisual dissonance issues in advance, which previously had to be judged retrospectively by human reviewers, thereby allowing a large amount of character content to be automatically generated with stable quality.

[0020] In addition, according to the present invention, since parameters are determined by relative deviation amounts corresponding to character attributes based on the individual vocal characteristics of the user as a reference point, different speech results are produced that naturally reflect the individual differences of each user even when different users possess the same character. This enhances both content diversity and user satisfaction in that the character voice is not fixed in a uniform and standardized manner, but maintains subtly different individuality for each speaker.

[0021] Furthermore, according to the present invention, new character attribute combinations not observed during the learning phase can be handled by interpolating and extrapolating potential representations in the attribute space. Given the nature of the character content industry, which generates countless new characters every year, this offers significant industrial applicability by providing scalability that can be applied immediately without requiring a separate learning process for each new character.

[0022] Furthermore, the present invention supports real-time streaming processing and multi-speaker / multi-character parallel processing, making it immediately applicable to various application fields requiring real-time interaction, such as live broadcasting, video conferencing, online games, the metaverse, and virtual and augmented reality content. This differentiates it from conventional offline methods that require post-processing, enabling a new form of real-time immersive content where users instantly become one with a character and converse. Brief explanation of the drawing

[0023] FIG. 1 is an overall conceptual diagram of an artificial intelligence-based customized character voice conversion and video generation system according to one embodiment of the present invention. FIG. 2 is a hardware block diagram of a voice conversion and video generation device according to one embodiment of the present invention. FIG. 3 is a software module configuration diagram according to one embodiment of the present invention. FIG. 4 is an overall flowchart of a character voice conversion and video generation method according to one embodiment of the present invention. FIG. 5 is a detailed flowchart of character attribute information acquisition and analysis according to one embodiment of the present invention. FIG. 6 is a detailed flowchart of the voice conversion parameter determination step according to one embodiment of the present invention. FIG. 7 is a detailed flowchart of the steps for combining a visual model and a target voice and generating a video according to one embodiment of the present invention. FIG. 8 is a data structure diagram of character attribute information according to one embodiment of the present invention. FIG. 9 is an example diagram of attribute information by various character types according to one embodiment of the present invention. FIG. 10 is a graph showing the inverse mapping relationship between the external size of a character and voice pitch and formant according to one embodiment of the present invention. FIG. 11 is a diagram of the principle for calculating the virtual star length and determining the formant shift coefficient according to one embodiment of the present invention. FIG. 12 is a diagram of the impulse response database configuration and convolution processing process by material according to one embodiment of the present invention. FIG. 13 is a diagram illustrating the principle of deriving the Helmholtz resonance frequency from the internal cavity structure of a character according to one embodiment of the present invention. FIG. 14 is a mapping table of acoustic texture imparting parameters by material category according to one embodiment of the present invention. FIG. 15 is a learning pipeline of an artificial intelligence-based speech conversion model according to one embodiment of the present invention. FIG. 16 is a network architecture of a conditional generation model according to one embodiment of the present invention. FIG. 17 is a diagram of the interpolation and extrapolation principles for a potential representation in an attribute space and a zero-shot character correspondence according to one embodiment of the present invention. FIG. 18 is a multimodal processing configuration diagram that utilizes visual features according to one embodiment of the present invention as additional inputs to an attribute encoder. FIG. 19 is a process diagram of speech timing and phoneme analysis and lip movement mapping of target voice data according to one embodiment of the present invention. FIG. 20 is an example diagram of an alternative utterance part by character type according to one embodiment of the present invention. FIG. 21 is a conceptual diagram of a finite element vibration model for physical vibration simulation of an alternative visual element according to one embodiment of the present invention. FIG. 22 is a flowchart of mechanical noise addition processing by driving mechanism type of a robot character according to one embodiment of the present invention. FIG. 23 is a process diagram for deriving voice conversion parameters from expected body parameters of a virtual race character according to one embodiment of the present invention. FIG. 24 is a diagram showing the additional processing of acoustic characteristics by species of an animal-type character according to one embodiment of the present invention. FIG. 25 is a flowchart showing the structure and feedback readjustment of an audiovisual matching determination model according to one embodiment of the present invention. FIG. 26 is a diagram of the principle for extracting user characteristics and determining relative parameters according to one embodiment of the present invention. FIG. 27 is a configuration diagram of a pipeline and latency management for real-time streaming processing according to an embodiment of the present invention. FIG. 28 is a diagram of a combination configuration within a single video frame and a multi-speaker / multi-character parallel processing according to an embodiment of the present invention. FIG. 29 is a flowchart of user emotion extraction and reflection in target voice data according to one embodiment of the present invention. FIG. 30 is a diagram of the UI and processing principle for adjusting the mixing ratio between the user individuality preservation strength and the character characteristic reflection strength according to one embodiment of the present invention. FIG. 31 is a flowchart of a conflict adjustment algorithm between user emotions and character personality traits according to one embodiment of the present invention. FIG. 32 illustrates a scene in which a user possesses a humanoid character and speaks as an example of use according to one embodiment of the present invention. FIG. 33 illustrates a scene in which a user possesses a robot character and speaks as an example of use according to an embodiment of the present invention. FIG. 34 illustrates a scene in which a user possesses an object or an abstract character and speaks, as an example of use according to an embodiment of the present invention. Specific details for implementing the invention

[0024] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings. However, the present invention is not limited or restricted by the following embodiments, and those skilled in the art should understand that all modifications, equivalents, and substitutions included within the spirit and scope of the present invention fall within the scope of the rights of the present invention. In assigning reference numerals to the components of each drawing, the same components are to have the same reference numeral as much as possible, even if they are shown in different drawings.

[0025] Referring to FIG. 1, an AI-based customized character voice conversion and video generation system according to one embodiment of the present invention is centered around a voice conversion and video generation device (100) that receives original voice data (30) from a user, obtains a visual model (10) and attribute information (20) of a target character, and finally outputs a video (60) in which the character speaks the content of the user's speech in the character's unique voice. The device (100) may be implemented as a personal computer, server, smartphone, tablet, VR / AR headset, game console, or a virtual instance on a cloud computing infrastructure, and may also be implemented in a form in which a user terminal and a server cooperate through a communication network to perform the method in a distributed manner.

[0026] Referring to FIG. 2, the voice conversion and video generation device (100) may include a processor (110), memory (120), an input / output unit (130), a communication unit (140), and a storage unit (150). The processor (110) is composed of a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), or a combination thereof, and the memory (120) includes RAM for storing instructions currently being executed and intermediate calculation results, and ROM for storing a boot loader, etc. The input / output unit (130) may include a microphone for collecting user speech, a display and speaker for playing the video (60), and a keyboard, touchscreen, controller, etc. for receiving user commands. The communication unit (140) transmits and receives data with an external server or another user terminal via a wired or wireless network, and the storage unit (150) stores a learned voice conversion model (400), a library of visual models (10) of characters, an impulse response database (46), and a pre-cached character-specific parameter set.

[0027] Referring to FIG. 3, the software module executed by the processor (110) is composed of an attribute analysis unit (200), a parameter determination unit (300), a voice conversion unit (400), a lip-sync processing unit (500), and a video generation unit (600). The attribute analysis unit (200) parses the attribute information (20) of the character to derive values ​​and categories for each attribute item. The parameter determination unit (300) maps the derived attribute values ​​to voice conversion parameters (40). The voice conversion unit (400) converts the original voice data (30) into target voice data (50) using the parameters (40) and an artificial intelligence-based voice conversion model. The lip-sync processing unit (500) generates changes in lip movements (530) or alternative visual elements (540) on the visual model (10) in accordance with the speech timing (510) and phoneme information (520) of the target voice data (50). The video generation unit (600) combines the lip-sync processing result and the target voice data (50) to render the final video (60).

[0028] Referring to FIG. 4, the overall flow of a method according to an embodiment of the present invention is described as follows: First, in step (a), the device (100) obtains a visual model (10) of the character, attribute information (20), and original voice data (30) of the user. Then, in step (b), the attribute analysis unit (200) and the parameter determination unit (300) analyze the attribute information (20) to determine voice conversion parameters (40). Then, in step (c), the voice conversion unit (400) converts the original voice data (30) into target voice data (50) using the parameters (40) and a learned voice conversion model. Finally, in step (d), the lip-sync processing unit (500) and the video generation unit (600) combine the visual model (10) and the target voice data (50) to generate a video (60).

[0029] The visual model (10) of the above character can be obtained in any form of visual representation, such as a 2D image, a 3D mesh, a neural radiance field (NeRF), a Gaussian splatting representation, or a live-action video clip. As an example, a doll photo taken by a user with a smartphone, a character file created by a game developer using a 3D modeling tool, a fantasy race illustration created by a text-to-image generation model, or an image of the user's own face taken with a camera can be used as the visual model (10).

[0030] Referring to FIG. 8, the attribute information (20) of the character has a hierarchical structure and includes sub-items such as whether it is biological or non-biological (21), species or object category (22), constituent material (23), external size and shape (24), age group (25), gender (26), personality traits (27), internal cavity structure information (28), and body parameter information (29). The attribute information (20) can be explicitly entered by a user through a UI, automatically extracted through a visual feature extraction unit (460) pre-learned from the visual model (10), parsed from metadata provided with the character, or derived through text description interpretation using a large-scale language model.

[0031] Referring to FIG. 9, examples of how the attribute information (20) is configured for various character types are illustrated. For example, in the case of a humanoid character, the biological / non-biological status (21) is set to 'biological', the species (22) to 'human', the material (23) to 'organic tissue', the size (24) to '170cm', the age group (25) to '30s', the gender (26) to 'male', and the personality (27) to 'calm'. In the case of a robot character, the biological / non-biological status (21) to 'non-biological', the category (22) to 'humanoid robot', the material (23) to 'metal-plastic composite', the size (24) to '3m large', and the internal cavity (28) to 'hollow chest'. In the case of the fairy character, (21) is set to ‘biological’, (22) to ‘virtual race: fairy’, (23) to ‘organic tissue’, (24) to ‘20cm small’, and body parameters (29) to ‘estimated spleen length 3cm, estimated lung capacity 10ml’.

[0032] Referring to FIG. 6, the parameter determination process of step (b) is described in detail. The parameter determination unit (300) maps each sub-item of the attribute information (20) to each component of the corresponding voice conversion parameter (40). Specifically, the external size (24) is inversely mapped to the pitch control value (41) and the formant control value (42), the constituent material (23) and the object category (22) are mapped to the resonance value (43), the mechanical sound filtering value (44), and the acoustic texture value (45), and the internal cavity structure (28) is mapped to the Helmholtz resonance frequency (47) and the resonance bandwidth (48). The age group (25), gender (26), and personality trait (27) are mapped to the fine-tuning value and speech style control value of the parameter.

[0033] With reference to FIGS. 10 and 11, the size-pitch inverse proportion mapping is described in detail. In voiced vibration vocalization mechanisms, including humans, it is acoustically established that the size of the vocal organ (vocal tract length) is inversely proportional to the resonance frequency. That is, the vocal tract length of an adult male is approximately 17 cm, that of a female is approximately 15 cm, and that of a child is approximately 12 cm; the longer the vocal tract, the lower the formant frequency is formed. The present invention extends and applies this physical principle to the size value (24) of a visual character. The parameter determination unit (300) calculates a virtual vocal tract length (L_v) from the external size (24) of the character, and, for example, may use the relationship L_v = k × H_c. Here, H_c is a value corresponding to the height of the character, and k is a coefficient by race / category.

[0034] Next, the parameter determination unit (300) calculates the ratio r = L_u / L_v between the calculated virtual vocal tract length L_v and the user's actual vocal tract length L_u, determines this ratio as a formant shift coefficient, and transmits it to the voice conversion unit (400). The user's actual vocal tract length L_u can be derived from the user's original voice data (30) through spectrum peak analysis or can use a value pre-stored in the user profile. The voice conversion unit (400) generates the target voice data (50) by shifting the formant frequency F_n of the original voice data (30) by F_n × r according to the formant shift coefficient r. For example, if the user's vocal cord length is 17cm and the character is a large robot with a height of 3m, L_v is calculated as 40cm, the formant shift coefficient r = 0.425, which lowers the formant of the original voice to about 42.5%, and as a result, a grand, low-pitched voice suitable for a large robot is generated. Conversely, if the character is a fairy with a height of 20cm, L_v is calculated as 3cm, r = 5.67, which raises the formant to about 5.67 times, and a very light, high-pitched fairy voice is generated.

[0035] The material-based acoustic texture imparting is described in detail with reference to FIGS. 12 and FIGS. 14. The storage unit (150) has a material-specific impulse response database (DB_46) pre-built. The DB_46 stores impulse response IR_m corresponding to each of a plurality of material categories, and each IR_m is obtained by (i) a method of playing an impulse signal on a real object composed of the corresponding material and measuring the response with a microphone, (ii) a method of numerically modeling the vibration and damping characteristics of the corresponding material through finite element analysis (FEM) simulation, or (iii) a method of extracting a spectrum envelope from a reference audio sample and inversely converting it into IR.

[0036] The above material categories may include the following as examples. Metal materials are characterized by bright initial decay and a long reverb tail, and a clear resonance peak formed in the 2–5 kHz band, expressed as IR; wood is characterized by mild mid-low frequency resonance and short reverberation, with a gentle peak around 500 Hz; ceramic and glass are characterized by a sharp resonance peak in the high frequency range and a very short decay time; plastic is characterized by relatively strong sound absorption, short reverberation, and a flat spectrum; rubber is characterized by rapid high frequency decay and dull low frequency; and organic tissue has soft resonance characteristics similar to the human vocal tract. The parameter determining unit (300) selects an IR_m corresponding to the constituent material (23) of the character from the DB_46, and the voice conversion unit (400) or post-processing convolution unit convolves the selected IR_m with the original voice data (30) or the intermediate output spectrogram of the voice conversion model to impart the material's inherent resonance characteristics.

[0037] Referring to FIG. 13, the derivation of the Helmholtz resonance parameter is explained in detail. When the character has a cavity inside, for example, a hollow robot chest, an empty doll body, or a jar-shaped object character, the cavity can be modeled as a Helmholtz Resonator. The Helmholtz resonance frequency f_H is calculated by the following formula: f_H = (c / 2ð) ? √(A / (V × L_n)), where c is the speed of sound, A is the cross-sectional area of ​​the cavity opening, V is the volume of the cavity, and L_n is the effective length of the opening. The parameter determination unit (300) derives the values ​​of A, V, and L_n from the internal cavity structure information (28) of the attribute information (20) and calculates f_H, and reflects the f_H in the resonance frequency component (47) of the voice conversion parameter (40). In addition, the resonance bandwidth (48) is derived from the common sound absorption characteristics to control whether the resonance is sharp or gentle. Through this processing, for example, the utterance of a jar character is given a low-frequency resonance characteristic of earthenware, and the utterance of a large robot with a hollow chest is given a chest cavity resonance, making it sound much more magnificent.

[0038] With reference to FIGS. 15 to 18, the structure and learning process of an artificial intelligence-based voice conversion model (400) will be described in detail. The voice conversion model (400) is composed of a Conditional Generative Model including a content encoder (410), an attribute encoder (420), and a decoder (430). The content encoder (410) receives the original voice data (30) of the user and produces a latent vector that encodes the speech content (31) and emotion information (32). The attribute encoder (420) receives the attribute information (20) of the character and embeds it into a conditional vector, and additionally receives the visual feature vector of the visual model (10) from the visual feature extraction unit (460) and combines it with the conditional vector. The decoder (430) combines the content latent vector and the condition vector to generate a Mel spectrogram of the target voice data (50), which is converted into a time domain waveform through a separate vocoder.

[0039] Referring to FIG. 15, the training data (440) of the voice conversion model (400) consists of voice data of various speakers and objects and pairs labeled with external attribute tags of each speaker and object. For example, each actual human voice sample is tagged with the person's height, age, gender, physique, etc., and machine speech samples recorded from robots of various sizes are tagged with the robot's size, material, and driving method, and the vocalizations of various animals are tagged with species, size, age, etc. Using the training data (440), the deep learning network (450) is trained to mimic the acoustic characteristics of the speech target according to the input attribute information. The learning loss function may be composed of a combination of (i) reconstruction loss (reconstructing the original speaker's utterance into the original speaker's attributes), (ii) attribute recursive loss (restoring the original when attribute A→B→A is transformed), and (iii) discriminator loss (adversarial learning).

[0040] Zero-shot character correspondence is explained with reference to FIG. 17. The attribute encoder (420) of the above-mentioned completed speech conversion model (400) is trained to learn continuously differentiable latent representations in the attribute space. Accordingly, even if a new combination of attributes not observed during the training phase, for example, 'a wooden giant with a height of 5m', is input, the attribute encoder (420) combines and extrapolates the learned latent representations of 'large size' and 'wooden material' to produce a corresponding condition vector. This is a very important characteristic in the practical environment of the content industry, where countless new characters that did not exist at the time of training are created during service, and it provides scalability that does not require retraining for every new character.

[0041] Lip-sync processing is described in detail with reference to FIG. 19. The lip-sync processing unit (500) decomposes the target voice data (50) into phoneme units and extracts the speech timing (510) of each phoneme (520). Subsequently, a viseme (visual speech unit) corresponding to each phoneme is queried, and the lip area (530) of the visual model (10) is animated frame by frame. In the case of a humanoid character, the blendshape coefficients of standard facial rigging are adjusted for each phoneme, and in the case of a 3D model, the joint angles of the jaw, tongue, and lip muscles may be directly adjusted.

[0042] Referring to FIG. 20, the processing of alternative speech parts of a non-humanoid character is described in detail. When the character does not have lips, the lip-sync processing unit (500) automatically identifies alternative visual elements (540) capable of representing speech on the visual model (10) or uses alternative elements specified by the user. Examples of alternative visual elements (540) include a speaker grille or light-emitting part of a robot, a lighting part of a doll, a vibrating part of an object, a waveform display part of a hologram, a shape-deforming part of an abstract shape, a surface light-emitting pattern of an egg-shaped character, and the opening and closing of an aperture of a robot child. Each alternative element changes over time according to the phoneme (520) and amplitude envelope (550) of the target voice data (50).

[0043] For example, if a robot's speaker grille is used as an alternative speech device, the grille's luminescence intensity changes in real time according to voice amplitude, allowing the user to visually recognize that the robot is currently speaking. If a doll has a light-emitting part, the luminescence color can be processed to subtly change depending on whether each phoneme is voiced or unvoiced. If an object character has a vibrating part, voice vibrations are converted into visual tremors and displayed. For abstract characters, the form can transform rhythmically according to the spectral characteristics of the voice.

[0044] Referring to FIG. 21, the physical vibration simulation of the alternate visual element is described. The lip-sync processing unit (500) maps the frequency spectrum of the target voice data (50) to the physical vibration mode (560, Vibration Mode) of the alternate visual element (540). A finite element vibration model (570) corresponding to the material (23) and shape (24) of the character is pre-constructed, and the vibration model pre-calculates and stores the unique vibration mode for each combination of material and shape. During rendering, each voice frequency component is converted into the excitation intensity of the corresponding vibration mode, and the surface displacement of the alternate visual element is calculated on a frame-by-frame basis as a superposition of these. For example, an alternate element in the shape of a thin metal plate is visually deformed into a disc vibration mode for low-frequency excitation and a multi-wave mode for high-frequency excitation, and this physically based rendering provides an immersive sensation in which the character's voice and visual vibration physically correspond.

[0045] Referring to FIG. 22, the reflection of mechanical drive characteristics of a robot character is explained. The parameter determining unit (300) identifies the type of drive mechanism (700) from the attribute information (20) of the character. The type of drive mechanism (700) includes, for example, one or a combination of an actuator (710), a servo motor (720), a gear train (730), and a hydraulic system (740). The parameter determining unit (300) includes additional signal information in the parameter (40) to add a mechanical noise component (750) corresponding to the identified type of drive mechanism (700) to the speech start and end sections of the target voice data (50). The mechanical noise component (750) may include motor howling (motor rotation sound), gear friction sound, hydraulic noise (hydraulic valve opening / closing sound), servo bouncing sound (fine vibration sound generated when determining servo position), etc.

[0046] For example, in the case of a large industrial robot character, a pressure rise noise from the hydraulic system (740) is briefly inserted just before the start of speech, and after the end of speech, the damped vibration sound of the actuator (710) is processed to fade out. In the case of a small toy robot, a fine popping sound from the servo motor (720) and a gear friction sound are subtly inserted between each syllable, giving an auditory sense that internal mechanical parts are moving whenever the robot speaks. This addition of mechanical noise provides a strong auditory realism that the resulting voice is actually the character speaking.

[0047] Referring to FIG. 23, the processing of non-existent virtual races is explained. When the character is a virtual race or fantasy creature that does not exist in reality, such as a fairy, dragon, alien, mythical creature, etc., the parameter determining unit (300) determines the voice conversion parameter (40) in the absence of actual voice data based on attribute information (29) regarding the character's body size, expected vocal tract structure, expected lung capacity, and expected vocal organ shape. For example, in the case of a dragon character, if attribute information is given such as a height of 30m, an expected vocal tract length of 150cm, an expected lung capacity of 500L, and an expected vocal organ as a 'large diaphragm including a double larynx,' the parameter determining unit (300) calculates a parameter including a very low f_H, an extremely low formant, and additionally a Growl component. As a result, the user's original voice is converted into a loud, low-pitched, threatening voice suitable for a dragon.

[0048] Animal-type character processing is described with reference to FIG. 24. When the character is an animal or an animal-type character, the parameter determination unit (300) adds at least one of a harmonic structure, a growl component, a chirp component, or a nasal resonance characteristic corresponding to the species information of the character to the voice conversion parameter (40). However, the intensity of the addition is limited to below a predetermined threshold so that the phoneme intelligibility of the user's original voice data (30) is maintained. For example, in the case of a cat-type character, a subtle purring harmonic and nasal resonance are added to the low frequency range, in the case of a bird-type character, a chirp component is subtly mixed in the high frequency range, and in the case of a bear-type character, a low-frequency growl component is added. These added components are applied within a range where the listener can still understand the language spoken by the user (e.g., a Korean sentence), creating an interesting auditory effect that is "like a cat speaking human language."

[0049] Cross-modal congruence verification is described with reference to FIG. 25. After step (c) and before step (d), the processor (110) calculates a cross-modal congruence score (810) between the visual model (10) of the character and the target voice data (50) through a pre-trained audiovisual congruence discrimination model (800). The discrimination model (800) calculates the congruence score (810) using the cosine similarity between the visual feature embeddings and the voice feature embeddings, or the output value of a discrimination network that receives both modalities together. If the score (810) is less than a predetermined threshold (e.g., 0.7), the processor (110) performs feedback readjustment (820) to cause the parameter determination unit (300) to readjust the voice conversion parameter (40), and then performs step (c) again. Through this closed-loop verification, audiovisual dissonance in the final output is automatically corrected within the system.

[0050] Normalization of user vocal tract characteristics is explained with reference to FIG. 26. In step (a) above, the device (100) extracts the user's vocal tract characteristics (33) from the user's original voice data (30). The vocal tract characteristics (33) include at least one of a fundamental frequency range F0, the position of formants F1 to F3, and a glottal closure constant. F0 is extracted using autocorrelation or a neural network-based pitch estimator, and F1 to F3 are extracted through Linear Predictive Coding (LPC) analysis or spectrum peak detection. The parameter determination unit (300) determines the voice conversion parameter (40) as a relative deviation amount corresponding to the character's attribute information (20), using the extracted vocal tract characteristics (33) as a reference point. That is, if User A is an adult male with an average F0 of 120Hz and User B is an adult female with an average F0 of 220Hz and possesses the same fairy character, the target F0 value of the fairy character is not absolutely fixed but is determined by a value increased by a certain proportion from each user's F0, so that the speech results of the two users maintain different personalities but are both naturally produced as voices suitable for the fairy.

[0051] Real-time streaming processing is described with reference to FIG. 27. Steps (a) through (d) can be performed in a streaming manner substantially simultaneously with the utterance of the user's original voice data (30). In this case, the parameter determination and conversion operations of steps (b) and (c) are optimized so that the delay time from the start of the user's utterance to the output of the corresponding frame of the video (60) is 300 milliseconds or less. Specifically, a set of character-specific parameters is pre-cached in the storage unit (150), and when a user selects a specific character during the service, the parameters (40) of that character are immediately loaded into the memory (120). The voice conversion unit (400) processes the original voice data (30) by dividing it into frames of 10 to 30 milliseconds and processing them in a streaming manner, and the conversion result of each frame is immediately transmitted to the lip-sync processing unit (500) and output together with the corresponding video frame. This real-time pipeline enables applications such as live broadcasting, VR avatar conversation, and character chatting within an online game.

[0052] Multi-speaker and multi-character processing is explained with reference to FIG. 28. In step (a), original voice data (30-1, 30-2, ..., 30-n) from multiple users and visual models (10-1, 10-2, ..., 10-n) of multiple characters are obtained, and steps (b) through (d) are performed in parallel according to the user-character correspondence (900). Each user-character pair is processed through a separate voice conversion pipeline, and finally, in step (d), a scene is created in which the multiple characters are placed together within a single video (60) frame and converse with each other. For example, when four users each possess a human, robot, fairy, or dragon character and conduct a real-time meeting, the speech of each user is converted into the voice of the corresponding character, and a video is created in which the four characters appear together on the screen and converse in real-time.

[0053] Emotion preservation is explained with reference to FIG. 29. In step (a) above, the device (100) extracts an emotional state (32) from the original voice data (30) of the user. The emotional state (32) is extracted through a pre-trained emotion recognition model and can be expressed as discrete emotion categories (joy, sadness, anger, surprise, fear, disgust, neutrality, etc.) or continuous emotion dimensions (valence, arousal, dominance). In step (c) above, the voice conversion unit (400) converts the extracted emotional state (32) by reflecting it in the intonation, timbre, speaking rate, and stress distribution of the target voice data (50). Accordingly, the user's original emotion is maintained and output as a voice that matches the characteristics of the character. For example, if a user speaks with the emotion of sadness, the characteristics of sadness—such as a low intonation, slow speech speed, and trembling at the end of words—are maintained even after being converted into the fairy character's voice, allowing the listener to naturally feel that the fairy is sad.

[0054] Referring to FIG. 30, the control of the user personality preservation strength is described. Step (c) can individually adjust the mixing ratio between the character characteristic reflection strength and the user personality preservation strength for each personality element (34) included in the user's original voice data (30). The personality element (34) includes at least one of an intonation pattern, speech rhythm, stress distribution, and dialect characteristics. The mixing ratio can be set by the user through a user interface (e.g., a slider for each personality element) or can be automatically determined based on the character's attribute information (20). For example, if the user sets it to 'maintain the Gyeongsang dialect but follow the characteristics of the fairy character's intonation,' the resulting voice has a bright, high intonation characteristic of a fairy, while the sentence ending processing maintains the characteristics of the Gyeongsang dialect.

[0055] Referring to FIG. 31, the adjustment of conflicts between emotions and personality traits is explained. When the user's emotional state (32) conflicts with the personality trait (27) defined in the character's attribute information (20), the voice conversion unit (400) modulates the emotional expression method to match the character's personality trait (27), while preserving the polarity (positive or negative) and intensity of the emotion. For example, if a user speaks with the emotion of 'anger' against a character whose personality trait is 'calm and restrained,' the voice is modulated to express restrained anger with a low, suppressed tone and slightly trembling endings, rather than expressing anger with an agitated high pitch and rapid speech speed. This provides a balance point that allows the user's emotion to be clearly conveyed to the listener while maintaining the character's persona.

[0056] FIG. 32 illustrates an example of use in which a user possesses a humanoid character. The user takes a picture or selects an image of a celebrity, historical figure, or virtual human character using their smartphone camera and loads the image into a visual model (10). When the user speaks into the microphone, the system converts the user's voice to match the characteristics of the person and generates a lip-synced video (60) in real time. This can be used for personal video production, educational content, social media posts, etc.

[0057] FIG. 33 illustrates an example of a user possessing a robot character. The user designates their robot avatar from a 3D game as a visual model (10) and inputs information regarding the robot's size, material, and driving mechanism as attribute information (20). The user's voice is converted into a large robot's majestic bass voice, and a video (60) is generated in which the robot's speaker grille glows in sync with speech. This can be utilized in online games, VR social platforms, marketing content utilizing robot characters, etc.

[0058] FIG. 34 illustrates an example of use in which a user possesses an object or an abstract character. The user designates an image of an object they like (e.g., pottery, a car, a tree) or an abstract shape (e.g., a glowing polyhedron, a color-changing organism) as a visual model (10). The system automatically extracts the material, size, shape, etc. of the object or converts the voice according to attribute information (20) obtained through user input, and lip-syncs by utilizing the surface waveform, glow, or deformation of the object as a speech part (540). This can be used in children's educational animations (content where objects speak), advertisements, artistic experiments, etc.

[0059] The embodiments described above are presented as examples to aid in understanding the present invention and are not limited thereto. Those skilled in the art to which the present invention pertains can make various modifications, combinations, and applications within the scope of the technical spirit of the present invention, and all such modifications, combinations, and applications should be interpreted as falling within the scope of the rights of the present invention. For example, the components individually described in the above embodiments may be combined and implemented together, and the features described in one embodiment may be combined with the features of another embodiment. Explanation of the symbols

[0060] 10: Visual model of the character 20: Character Attribute Information 21: Biological / Non-biological status 22: Species or Object Category 23: Composition Material 24: External Size and Shape 25: Age group 26: Gender 27: Personality Traits 28: Internal Joint Structure Information 29: Body Parameter Information 30: User's original voice data 31: Utterance content 32: Emotional State 33: User Saint Characteristics 34: Personality Elements 40: Speech conversion parameters 41: Pitch adjustment value 42: Formant adjustment value 43: Resonance value 44: Machine sound filtering value 45: Acoustic texture value 46: Impulse Response Database 47: Helmholtz resonance frequency 48: Resonance Bandwidth 50: Target voice data 60: Video 100: Voice conversion and video generation device 110: Processor 120: Memory 130: Input / Output Section 140: Communications Department 150: Storage section 200: Attribute Analysis Department 300: Parameter determination unit 400: Speech conversion unit (speech conversion model) 410: Content Encoder 420: Attribute Encoder 430: Decoder 440: Training data 450: Deep Learning Network 460: Visual Feature Extraction Unit 500: Lip-sync processing unit 510: Ignition Timing 520: Phoneme Information 530: Lip movements 540: Alternative visual elements 550: Amplitude envelope 560: Physical vibration mode 570: Finite element vibration model 600: Video generation section 700: Driving mechanism 710: Actuator 720: Servo motor 730: Gear Train 740: Hydraulic System 750: Mechanical noise component 800: Audiovisual Coherence Determination Model 810: Cross-modal consistency score 820: Feedback readjustment 900: User-Character Correspondence

Claims

Claim 1 A method for customized character voice conversion and video generation using artificial intelligence, comprising the following steps performed by a computing device: (a) acquiring a visual model of a target character, attribute information of the character, and original voice data of a user; (b) determining voice conversion parameters by analyzing the attribute information of the character; (c) converting the original voice data of the user into target voice data corresponding to the characteristics of the character using the determined voice conversion parameters and an artificial intelligence-based voice conversion model; and (d) generating a video by combining the acquired visual model of the character and the target voice data. Claim 2 A method according to claim 1, characterized in that the attribute information of the character includes at least one of whether it is biological or non-biological, species or object category, constituent material, external size and shape, age group, gender, or personality trait. Claim 3 A method according to claim 1, wherein step (d) comprises the step of generating the video by lip-syncing the visual model of the character so that it moves its lips or speech part in accordance with the speech timing and pronunciation structure (Phoneme) of the target voice data. Claim 4 A method according to claim 1, wherein the voice conversion parameter of step (b) comprises a value that adjusts the pitch and formant of the voice inversely proportional to the external size of the character, and a value that applies reverberation, mechanical sound filtering, or acoustic texture effects according to the constituent material or object category characteristics of the character. Claim 5 A method according to claim 1, wherein the artificial intelligence-based speech conversion model of step (c) is a model trained through a deep learning network to simulate the acoustic characteristics of a speech target according to input attribute information, based on training data in which voice data of various speakers and objects and external attribute tags of the speakers and objects are labeled in advance. Claim 6 A method according to claim 1, further comprising the step of extracting an emotional state from the original voice data of the user in step (a), and in step (c), converting the extracted emotional state of the user by reflecting it in the intonation and timbre of the target voice data so as to output a voice that matches the characteristics of the character while maintaining the user's original emotion. Claim 7 In claim 1, the step (b) comprises: a step of calculating a virtual vocal tract length from the external size of the character; and a step of determining a formant shift coefficient based on the ratio between the calculated virtual vocal tract length and the user's actual vocal tract length, and in the step (c), a method characterized by generating the target voice data by shifting the formant frequency of the user's original voice data according to the determined formant shift coefficient. Claim 8 In claim 4, the imparting of an acoustic texture effect according to the constituent material of the character comprises: a step of selecting an impulse response corresponding to the constituent material of the character from an impulse response database that has been measured or modeled in advance corresponding to each of a plurality of material categories; and a step of convolving the selected impulse response with the original voice data or the intermediate output of the voice conversion model to impart the inherent resonance characteristics of the material, wherein the plurality of material categories include at least one of metal, wood, ceramic, glass, plastic, rubber, or organic tissue. Claim 9 In claim 4, the attribute information of the character further includes cavity structure information within the character, and the step (b) is characterized by deriving the volume and shape of the cavity from the cavity structure information to determine the Helmholtz resonance frequency and resonance bandwidth, and reflecting the determined resonance frequency and resonance bandwidth in the voice conversion parameters. Claim 10 A method according to claim 1, wherein, if the character is a robot or a mechanical device, step (b) comprises: identifying the type of driving mechanism from the attribute information of the character; and adding a mechanical noise component corresponding to the identified type of driving mechanism to the speech start and end intervals of the target voice data, wherein the type of driving mechanism includes at least one of an actuator, a servo motor, a gear train, or a hydraulic system, and the mechanical noise component includes at least one of motor howling, gear friction sound, hydraulic noise, or servo bouncing sound. Claim 11 A method according to claim 1, wherein if the character is a virtual race or fantasy creature that does not exist in reality, the step (b) is characterized by determining the voice conversion parameters in the absence of actual voice data based on attribute information regarding the character's body size, expected vocal tract structure, expected lung capacity, and expected vocal organ shape. Claim 12 In claim 1, if the character is an animal or an animal-type character, the method is characterized in that step (b) adds at least one of a harmonic structure, a growl component, a chirp component, or a nasal resonance characteristic corresponding to the species information of the character to the voice conversion parameter, while limiting the intensity of the addition to a threshold value or lower so as to maintain the phoneme intelligibility of the user's original voice data. Claim 13 A method according to claim 1, further comprising: a step of calculating a cross-modal congruence score between the visual model of the character and the target voice data through a pre-trained audiovisual congruence discrimination model after step (c) and before step (d); and a feedback step of readjusting the voice conversion parameters of step (b) and re-performing step (c) if the congruence score is less than a threshold value. Claim 14 A method according to claim 1, wherein step (a) further comprises the step of extracting vocal tract characteristics of the user from the original voice data of the user, wherein the vocal tract characteristics of the user include at least one of a fundamental frequency (F0) range, the position of formants F1 to F3, or a glottal closure constant, and wherein the voice conversion parameter of step (b) is determined as a relative deviation amount corresponding to the attribute information of the character using the extracted vocal tract characteristics of the user as a reference point. Claim 15 A method according to claim 3, wherein, in the case where the character does not have lips, the lip-sync processing identifies an alternative visual element representing speech on a visual model of the character and processes the alternative visual element to change according to the phoneme and amplitude envelope of the target voice data, wherein the alternative visual element includes at least one of a speaker grille or light-emitting part of a robot, a lighting part of a doll, a vibrating part of an object, a waveform display part of a hologram, or a shape-changing part of an abstract shape. Claim 16 A method according to claim 15, wherein the change in the alternative visual element is rendered based on a finite element vibration model corresponding to the material and shape of the character by mapping the frequency spectrum of the target voice data to the physical vibration mode of the alternative visual element. Claim 17 A method according to claim 1, wherein steps (a) to (d) are performed in a streaming manner substantially simultaneously with the utterance of the user's original voice data, and wherein the parameter determination and conversion operations of steps (b) and (c) are performed using a pre-cached character-specific parameter set such that the delay time from the start of the user's utterance to the output of the corresponding frame of the video is 300 milliseconds or less. Claim 18 A method according to claim 1, wherein in step (a), original voice data from a plurality of users and visual models of a plurality of characters are obtained, steps (b) to (d) are each performed in parallel according to the user-character correspondence, and in step (d), a scene is created in which the plurality of characters are placed together within a single video frame and interact with each other. Claim 19 In claim 5, the AI-based voice conversion model comprises a conditional generative model including: a content encoder that encodes linguistic content and emotion from the user's original voice data; an attribute encoder that embeds attribute information of the character into a conditioning vector; and a decoder that combines the output of the content encoder and the output of the attribute encoder to generate a mel-spectrogram of the target voice data, wherein the attribute encoder is configured to receive visual features extracted from the character's visual model as additional input. Claim 20 A method according to claim 5, wherein the artificial intelligence-based speech conversion model is trained to learn a latent representation that is continuously differentiable in an attribute space, so as to determine the speech conversion parameters by interpolating and extrapolating the learned attribute-acoustic feature mapping even for new character attribute combinations not observed during the learning phase. Claim 21 In claim 6, the above step (c) is configured to individually adjust the mixing ratio between the character characteristic reflection strength and the user personality preservation strength for each personality element included in the user's original voice data, wherein the personality element includes at least one of an intonation pattern, speech rhythm, stress distribution, or dialect characteristic, and the mixing ratio is set by the user through a user interface or automatically determined based on the attribute information of the character. Claim 22 A method according to claim 6, wherein when the emotional state of the user conflicts with the personality traits defined in the attribute information of the character, the method of expressing the user's emotion is modified to conform to the personality traits of the character, while preserving the polarity (positive or negative) and intensity of the emotion. Claim 23 A device for customized character voice conversion and video generation using artificial intelligence, comprising: a memory for storing commands; and a processor for executing commands, wherein the processor acquires a visual model of a target character, attribute information of the character, and original voice data of a user, converts the original voice data into target voice data through an artificial intelligence-based voice conversion model based on voice conversion parameters determined by analyzing the attribute information of the character, and then generates a video by combining the acquired visual model of the character and the target voice data.