Image generation methods, apparatus, storage media and electronic devices
By generating images from user voice data, the problem of existing technologies being unable to reflect users' psychological state is solved, enabling visualization of psychological state and health assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QINGDAO HAIER TECH
- Filing Date
- 2022-06-29
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies cannot generate images that correspond to the user's psychological state, and therefore cannot reflect the user's psychological state.
By sending preset text to smart devices, the system parses keywords and voice feature information from user voice data, generates image templates using the voice feature information and keywords, and adjusts image features to reflect the user's psychological state.
It enables the generation of images corresponding to a user's psychological state based on the user's voice characteristics, reflecting changes in the user's psychological state and providing mental health assessments and prompts.
Smart Images

Figure CN115269906B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home technology, and more specifically, to an image generation method, apparatus, storage medium, and electronic device. Background Technology
[0002] Currently, with the rapid development of science and technology, digital image processing technology in the field of image processing is also advancing rapidly. For example, it is now possible to use infrared image dynamic simulation software to organize parameters such as temperature field and temperature reflectivity of an object at various angles into floating-point format DDS (Direct Draw Surface) textures, and to generate target images based on the coordinate mapping relationship between these two-dimensional textures surrounding the object at different angles and the vertices of the three-dimensional model. However, this target image generated from the acquired infrared data of the object can only reflect the appearance of the object and cannot reflect the psychological state of the object.
[0003] Therefore, in related technologies, there is a problem of how to generate images that correspond to the user's psychological state.
[0004] No effective solution has yet been proposed for the problem of how to generate images that correspond to the user's psychological state in related technologies. Summary of the Invention
[0005] This application provides an image generation method, apparatus, storage medium, and electronic device to at least solve the problem in the related art of how to generate images that correspond to the user's psychological state.
[0006] According to one embodiment of this application, an image generation method is provided, comprising: sending preset text to a smart device; parsing voice data uploaded by the smart device to obtain keywords of the voice data and voice feature information corresponding to the voice data, wherein the voice data is voice data generated when a first object reads the preset text; and generating a target image corresponding to the first object based on the voice feature information and an image template corresponding to the keywords.
[0007] In an exemplary embodiment, generating a target image corresponding to the first object based on the speech feature information and the image template corresponding to the keyword includes: performing feature recognition on the keywords of the speech data to obtain the word category corresponding to the keyword; obtaining an image template pre-set for the word category and an image feature adjustment value corresponding to the speech feature information; and generating a target image based on the image feature adjustment value and the image template pre-set for the word category.
[0008] In an exemplary embodiment, obtaining the image feature adjustment value corresponding to the speech feature information includes: obtaining the speech volume in the speech feature information, and determining the first difference between the speech volume and the first preset value as the image feature adjustment value when the speech volume is greater than a first preset value and less than a second preset value; generating a target image based on the image feature adjustment value and an image template preset for the word category includes: adjusting the size of the initial image in the image template according to the image feature adjustment value to obtain the target image.
[0009] In an exemplary embodiment, obtaining the image feature adjustment value corresponding to the speech feature information includes: obtaining the speech jitter frequency in the speech feature information, and determining the second difference between the speech jitter frequency and the third preset value as the image feature adjustment value when it is determined that the speech jitter frequency is greater than a third preset value; generating a target image based on the image feature adjustment value and an image template preset for the word category includes: adjusting the width of the outline lines of the initial image in the image template according to the image feature adjustment value to obtain the target image.
[0010] In an exemplary embodiment, obtaining the image feature adjustment value corresponding to the speech feature information includes: obtaining the speech rate of the speech feature information within a preset time period, and determining the third difference between the speech rate and the fourth preset value as the image feature adjustment value if the speech rate is determined to be greater than a fourth preset value; generating a target image based on the image feature adjustment value and an image template preset for the word category includes: adjusting the tilt angle of the contour lines of the initial image in the image template according to the image feature adjustment value to obtain the target image.
[0011] In an exemplary embodiment, generating a target image corresponding to the first object based on the speech feature information and the image template corresponding to the keyword includes: obtaining the number of different pitch values from the speech feature information; if the number of different pitch values is greater than a fifth preset value, determining the hue type of the initial image in the image template as a first color system to obtain the target image.
[0012] In one exemplary embodiment, the method further includes: acquiring multiple voice feature dimensions corresponding to the first voice feature information, wherein the multiple voice feature dimensions include at least one of the following: voice volume dimension, voice jitter frequency dimension, and speech rate dimension; acquiring multiple psychological stress values of the first voice feature information under the multiple voice feature dimensions; determining abnormal stress values from the multiple psychological stress values; and, if the number of abnormal stress values is greater than a sixth preset value, sending a prompt message to the first object to indicate that the abnormal stress value exists.
[0013] According to another embodiment of this application, an image generation apparatus is also provided, comprising: a sending module for sending preset text to a smart device; a parsing module for parsing voice data uploaded by the smart device to obtain keywords of the voice data and voice feature information corresponding to the voice data, wherein the voice data is voice data generated when a first object reads the preset text; and a determining module for generating a target image corresponding to the first object based on the voice feature information and an image template corresponding to the keywords.
[0014] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, which is configured to execute the above-described image generation method at runtime.
[0015] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the image generation method described above through the computer program.
[0016] In this embodiment, a preset text is sent to a smart device; the voice data uploaded by the smart device is parsed to obtain keywords of the voice data and voice feature information corresponding to the voice data, wherein the voice data is voice data generated when a first object reads the preset text; a first image corresponding to the first object is generated based on the voice feature information and the image template corresponding to the keywords; by adopting the above technical solution, the problem of how to generate an image corresponding to the user's psychological state is solved. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the hardware environment for an image generation method according to an embodiment of this application;
[0020] Figure 2 This is a flowchart of an image generation method according to an embodiment of this application;
[0021] Figure 3 This is a schematic diagram of an image generation method according to an embodiment of this application;
[0022] Figure 4 This is a structural block diagram (a) of an image generation apparatus according to an embodiment of this application;
[0023] Figure 5 This is a structural block diagram (II) of an image generation apparatus according to an embodiment of this application. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] According to one aspect of the embodiments of this application, an image generation method is provided. This image generation method is widely used in whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligencehouse ecosystems. Optionally, in this embodiment, the above-mentioned image generation method can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.
[0027] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.
[0028] This embodiment provides an image generation method applied to the aforementioned computer terminal. Figure 2 This is a flowchart of an image generation method according to an embodiment of this application, the process including the following steps:
[0029] Step S202: Send a preset text to the smart device;
[0030] Step S204: Analyze the voice data uploaded by the smart device to obtain the keywords of the voice data and the voice feature information corresponding to the voice data. The voice data is the voice data generated when the first object reads the preset text.
[0031] It should be noted that the aforementioned speech feature information may include, but is not limited to, speech volume, speech rate, speech jitter frequency, speech timbre, and speech tone, and this application does not impose any restrictions on these aspects.
[0032] Step S206: Generate the target image corresponding to the first object based on the speech feature information and the image template corresponding to the keyword.
[0033] The image templates corresponding to the keywords can be pre-set or generated in real time. For example, for real-time generated image templates, the scene can be generated based on keywords representing nouns, i.e., the image template itself. Alternatively, the positional relationship between scenes can be generated based on keywords representing adjectives.
[0034] It should be noted that the above image template can be a template with only a black and white background, or it can include a color template; this application does not impose any restrictions on this.
[0035] Through the above steps, a preset text is sent to a smart device; the voice data uploaded by the smart device is parsed to obtain the keywords of the voice data and the corresponding voice feature information, wherein the voice data is the voice data generated when the first object reads the preset text; and a target image corresponding to the first object is generated based on the voice feature information and the image template corresponding to the keywords, thus solving the problem in related technologies of how to generate an image corresponding to the user's psychological state.
[0036] In an exemplary embodiment, to better understand how step S206 generates the target image corresponding to the first object based on the speech feature information and the image template corresponding to the keyword, the following scheme can be used: perform feature recognition on the keywords of the speech data to obtain the word category corresponding to the keyword; obtain the image template pre-set for the word category and the image feature adjustment value corresponding to the speech feature information; generate the target image based on the image feature adjustment value and the image template pre-set for the word category.
[0037] It should be noted that the above-mentioned word categories can include nouns, verbs, adjectives, numerals, measure words, pronouns, adverbs, prepositions, conjunctions, auxiliary words, interjections, onomatopoeia, etc., but are not limited to these. Furthermore, different word categories correspond to different image processing methods. For example, nouns can directly generate the corresponding scenery, while conjunctions, adverbs, auxiliary words, etc., which have no actual meaning, can be used to adjust the positions of scenery objects, etc.
[0038] In one exemplary embodiment, several schemes are further proposed for obtaining image feature adjustment values corresponding to the speech feature information, and for generating a target image based on the image feature adjustment values and an image template pre-set for the word category, specifically including:
[0039] Solution 1: Obtaining the image feature adjustment value corresponding to the speech feature information includes: obtaining the speech volume in the speech feature information, and when the speech volume is greater than a first preset value and less than a second preset value, determining the first difference between the speech volume and the first preset value as the image feature adjustment value; generating a target image based on the image feature adjustment value and an image template preset for the word category includes: adjusting the size of the initial image in the image template according to the image feature adjustment value to obtain the target image.
[0040] The above embodiments not only enable the adjustment of image size according to the user's voice volume, but also realize the scheme of generating target images based on voice volume. That is, the target image corresponding to the user's psychological state can be generated based on the changing voice volume of the user under different psychological states.
[0041] Option 2: Obtaining the image feature adjustment value corresponding to the speech feature information includes: obtaining the speech jitter frequency in the speech feature information, and determining the second difference between the speech jitter frequency and the third preset value as the image feature adjustment value when the speech jitter frequency is determined to be greater than a third preset value; generating a target image based on the image feature adjustment value and an image template preset for the word category includes: adjusting the width of the outline lines of the initial image in the image template according to the image feature adjustment value to obtain the target image.
[0042] It is understood that in this embodiment, the width of the outline lines of the initial image in the image template can also be adjusted according to the intensity of the voice jitter frequency over a period of time.
[0043] Furthermore, the process of adjusting the outline line width of the initial image in the image template according to the second difference between the voice jitter frequency and the third preset value to obtain the target image may include: adjusting the outline line width of the initial image according to the line width adjustment ratio corresponding to the second difference to obtain the target image.
[0044] If the voice jitter frequency is determined to be less than a third preset value, then the contour line width corresponding to the third preset value is directly used as the contour line width of the initial image to obtain the target image.
[0045] The above embodiments realize a scheme to adjust the width of the outline lines of the initial image in the image template according to the frequency of voice jitter. That is, a target image corresponding to the user's psychological state can be generated according to the frequency of voice jitter of the user in different psychological states.
[0046] Solution 3: Obtaining the image feature adjustment value corresponding to the speech feature information includes: obtaining the speech rate of the speech feature information within a preset time period, and determining the third difference between the speech rate and the fourth preset value as the image feature adjustment value when the speech rate is determined to be greater than a fourth preset value; generating a target image based on the image feature adjustment value and an image template preset for the word category includes: adjusting the tilt angle of the contour lines of the initial image in the image template according to the image feature adjustment value to obtain the target image.
[0047] It should be noted that the tilt angle of the outline of the initial image can be understood as the angle relative to any perpendicular line of the image template. In the case that the initial image is a mountain peak, the tilt angle of the outline of the mountain peak can be understood as the tilt angle of the peak tip relative to the perpendicular line of the mountain peak.
[0048] Furthermore, the process of adjusting the tilt angle of the contour lines according to the third difference between the speech rate and the fourth preset value to obtain the target image may include: adjusting the tilt angle of the contour lines according to the tilt angle adjustment ratio corresponding to the third difference to obtain the target image.
[0049] If the speech rate is determined to be less than the fourth preset value, then the line tilt angle corresponding to the fourth preset value is used as the contour line tilt angle of the initial image.
[0050] The above embodiments realize a scheme to adjust the tilt angle of the outline lines of the initial image in the image template according to the speech rate. That is, a target image corresponding to the user's psychological state can be generated according to the user's speech rate in different psychological states.
[0051] In one exemplary embodiment, a technical solution is further proposed to generate a target image corresponding to the first object based on the speech feature information and the image template corresponding to the keyword. Specifically, it includes: obtaining the number of different pitch values from the speech feature information; and when the number of different pitch values is greater than a fifth preset value, determining the hue type of the initial image in the image template as a first color system to obtain the target image.
[0052] It is understandable that the first color system mentioned above can be a warm color system, such as red, orange, yellow, etc.
[0053] Optionally, in one embodiment, if the number of different tone values is less than a fifth preset value, the hue type of the image template is determined to be a second color system. The second color system can be a cool color system, such as blue, green, etc.
[0054] In one exemplary embodiment, the method further includes: acquiring multiple voice feature dimensions corresponding to the first voice feature information, wherein the multiple voice feature dimensions include at least one of the following: voice volume dimension, voice jitter frequency dimension, and speech rate dimension; acquiring multiple psychological stress values of the first voice feature information under the multiple voice feature dimensions; determining abnormal stress values from the multiple psychological stress values; and, if the number of abnormal stress values is greater than a sixth preset value, sending a prompt message to the first object to indicate that the abnormal stress value exists.
[0055] Among them, the sixth preset value mentioned above can be 0, for example. If the number of abnormal pressure values is determined to be 1, it can be determined that the first object has an abnormal pressure value under a certain feature dimension.
[0056] Furthermore, when the number of abnormal stress values exceeds a sixth preset value, sending a prompt message to the first object may further include: acquiring a first number of the aforementioned multiple psychological stress values and a second number of abnormal stress values; if the ratio of the second number to the first number is determined to be greater than a preset threshold, then a prompt message is sent to the first object to indicate that the first object has the abnormal stress value, and the first object may also be prompted to take stress reduction measures; wherein, the ratio of the second number to the first number indicates the degree of abnormal stress experienced by the first object, and is directly proportional to the degree of stress reduction achieved by the stress reduction measures.
[0057] In other embodiments, the following scheme is also proposed, specifically including: determining the voice volume of the first voice feature information in the voice volume dimension, the voice jitter frequency of the first voice feature information in the voice jitter dimension, and the voice speed of the first voice feature information in the voice speed dimension; and determining a first dimension threshold in the voice volume dimension, a second dimension threshold in the voice jitter dimension, and a third dimension threshold in the voice speed dimension; and determining that the first voice feature information has an abnormal pressure value in the voice speed dimension when the voice volume exceeds the first dimension threshold, or the voice jitter frequency exceeds the second dimension threshold, or the voice speed exceeds the third dimension threshold.
[0058] The method in this embodiment can be used in the most familiar and relaxing home environment for users. It scores and presents the user's mental health through emotion recognition and drawing creation.
[0059] To better understand the process of the above image generation method, the implementation flow of the above image generation method will be described below in conjunction with optional embodiments, but this is not intended to limit the technical solution of the embodiments of this application.
[0060] This embodiment provides an image generation method. Figure 3 This is a schematic diagram of an image generation method according to an embodiment of this application, such as... Figure 3 As shown, the specific steps are as follows:
[0061] Step S301: The user opens the "Emotion Drawing" program embedded in the smart screen by voice or click.
[0062] Step S302: A text describing a scene or image is displayed on the screen. For example: "The sun sets beyond the mountains, the Yellow River flows into the sea."
[0063] Step S303: The user reads the text aloud to the screen in a relaxed voice and tone, sentence by sentence.
[0064] Step S304: Analyze the text read aloud by the user. Match the entity words that can be drawn and the descriptive adjectives, such as: white, sun, mountain. Extract acoustic signal features such as volume (equivalent to the above-mentioned voice volume), voice wave jitter (equivalent to the above-mentioned voice jitter frequency), and speech rate from the user's voice.
[0065] Step S305: Analyze the user's emotions from the user's spoken voice.
[0066] As shown in Table 1 below, user emotions can be simply divided into two categories: positive and negative. Each category includes several subcategories of emotions, and each subcategory corresponds to N different emotional levels, where 1 represents the mildest and N represents the most severe.
[0067] Table 1
[0068]
[0069]
[0070] Step S306: Based on the entity words and descriptive words parsed in step S304, acoustic signal features, and user emotions parsed in step S305, create a drawing. The objects in the drawing are generated using entity words, the positional relationships between objects are generated using descriptive words, the size of the object outlines corresponds to the user's voice volume (the louder the voice, the larger the object), the smoothness of the object outlines corresponds to the degree of shaking in the user's voice (the more shaking the user's voice, the more shaking and thinner the outline), and the roundness of the object outlines corresponds to the user's speaking speed (the faster the user speaks, the steeper the mountain).
[0071] Positive emotions are filled with warm colors, while negative emotions are filled with cool colors. Each emotional subcategory corresponds to a different color brightness, and N emotional intensities correspond to N saturation levels.
[0072] Step S307: Score the user's mental health based on the drawing results.
[0073] Step S308: If the user's score is lower than the health threshold, notify the user and their family in a timely manner and provide methods to alleviate their emotions.
[0074] Through the above embodiments, a method for assessing mental health by combining voice emotion, acoustic signal features, and drawing is proposed. Specifically, acoustic signal features can be used to generate drawing outlines, color can be used to distinguish emotion categories, and color saturation can be used to distinguish emotion levels. Furthermore, users can conduct mental health assessments at home without leaving their homes, eliminating the need to consult a professional psychologist. In terms of usage, users only need to interact with a smart screen device to complete the assessment, protecting their personal privacy. Moreover, this method, using an entertaining and fun approach in a familiar home environment, allows users to relax more easily and obtain more accurate assessment results.
[0075] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0076] Figure 4 This is a structural block diagram (a) of an image generation apparatus according to an embodiment of this application. Figure 4 As shown, it includes:
[0077] Sending module 42 is used to send preset text to smart devices;
[0078] The parsing module 44 is used to parse the voice data uploaded by the smart device to obtain the keywords of the voice data and the voice feature information corresponding to the voice data. The voice data is the voice data generated when the first object reads the preset text.
[0079] It should be noted that the aforementioned speech feature information may include, but is not limited to, speech volume, speech rate, speech jitter frequency, speech timbre, and speech tone, and this application does not impose any restrictions on these aspects.
[0080] The determining module 46 is used to generate a target image corresponding to the first object based on the speech feature information and the image template corresponding to the keyword.
[0081] The image templates corresponding to the keywords can be pre-set or generated in real time. For example, for real-time generated image templates, the scene can be generated based on keywords representing nouns, i.e., the image template itself. Alternatively, the positional relationship between scenes can be generated based on keywords representing adjectives.
[0082] It should be noted that the above image template can be a template with only a black and white background, or it can include a color template; this application does not impose any restrictions on this.
[0083] The above-described device sends a preset text to a smart device; parses the voice data uploaded by the smart device to obtain keywords and corresponding voice feature information of the voice data, wherein the voice data is the voice data generated when a first object reads the preset text; and generates a target image corresponding to the first object based on the voice feature information and the image template corresponding to the keywords, thus solving the problem in related technologies of how to generate images corresponding to the user's psychological state.
[0084] Optionally, in an exemplary embodiment, the determining module 46 is further configured to perform feature recognition on the keywords of the speech data to obtain the word category corresponding to the keywords; obtain an image template pre-set for the word category and an image feature adjustment value corresponding to the speech feature information; and generate a target image based on the image feature adjustment value and the image template pre-set for the word category.
[0085] It should be noted that the above-mentioned word categories can include nouns, verbs, adjectives, numerals, measure words, pronouns, adverbs, prepositions, conjunctions, auxiliary words, interjections, onomatopoeia, etc., but are not limited to these. Furthermore, different word categories correspond to different image processing methods. For example, nouns can directly generate the corresponding scenery, while conjunctions, adverbs, auxiliary words, etc., which have no actual meaning, can be used to adjust the positions of scenery objects, etc.
[0086] Optionally, in an exemplary embodiment, the determining module 46 is further configured to obtain an image feature adjustment value corresponding to the speech feature information, including: obtaining the speech volume in the speech feature information, and determining a first difference between the speech volume and the first preset value as the image feature adjustment value when the speech volume is greater than a first preset value and less than a second preset value; generating a target image based on the image feature adjustment value and an image template preset for the word category, including: adjusting the size of the initial image in the image template according to the image feature adjustment value to obtain the target image.
[0087] The above embodiments not only enable the adjustment of image size according to the user's voice volume, but also realize the scheme of generating target images based on voice volume. That is, the target image corresponding to the user's psychological state can be generated based on the changing voice volume of the user under different psychological states.
[0088] Optionally, in an exemplary embodiment, the determining module 46 is further configured to acquire the speech jitter frequency in the speech feature information, and, if the speech jitter frequency is determined to be greater than a third preset value, determine the second difference between the speech jitter frequency and the third preset value as the image feature adjustment value; generate a target image based on the image feature adjustment value and an image template preset for the word category, including: adjusting the outline line width of the initial image in the image template according to the image feature adjustment value to obtain the target image.
[0089] It is understood that in this embodiment, the width of the outline lines of the initial image in the image template can also be adjusted according to the intensity of the voice jitter frequency over a period of time.
[0090] Furthermore, the process of adjusting the outline line width of the initial image in the image template according to the second difference between the voice jitter frequency and the third preset value to obtain the target image may include: adjusting the outline line width of the initial image according to the line width adjustment ratio corresponding to the second difference to obtain the target image.
[0091] If the voice jitter frequency is determined to be less than a third preset value, then the contour line width corresponding to the third preset value is directly used as the contour line width of the initial image to obtain the target image.
[0092] The above embodiments realize a scheme to adjust the width of the outline lines of the initial image in the image template according to the frequency of voice jitter. That is, a target image corresponding to the user's psychological state can be generated according to the frequency of voice jitter of the user in different psychological states.
[0093] Optionally, in an exemplary embodiment, the determining module 46 is further configured to acquire the speech rate of the speech feature information within a preset time period, and, if the speech rate is determined to be greater than a fourth preset value, determine the third difference between the speech rate and the fourth preset value as the image feature adjustment value; generate a target image based on the image feature adjustment value and an image template preset for the word category, including: adjusting the tilt angle of the contour lines of the initial image in the image template according to the image feature adjustment value to obtain the target image.
[0094] It should be noted that the tilt angle of the outline of the initial image can be understood as the angle relative to any perpendicular line of the image template. In the case that the initial image is a mountain peak, the tilt angle of the outline of the mountain peak can be understood as the tilt angle of the peak tip relative to the perpendicular line of the mountain peak.
[0095] Furthermore, the process of adjusting the tilt angle of the contour lines according to the third difference between the speech rate and the fourth preset value to obtain the target image may include: adjusting the tilt angle of the contour lines according to the tilt angle adjustment ratio corresponding to the third difference to obtain the target image.
[0096] If the speech rate is determined to be less than the fourth preset value, then the line tilt angle corresponding to the fourth preset value is used as the contour line tilt angle of the initial image.
[0097] The above embodiments realize a scheme to adjust the tilt angle of the outline lines of the initial image in the image template according to the speech rate. That is, a target image corresponding to the user's psychological state can be generated according to the user's speech rate in different psychological states.
[0098] Optionally, in an exemplary embodiment, the determining module 46 is further configured to obtain the number of different pitch values from the speech feature information; if the number of different pitch values is greater than a fifth preset value, determine the hue type of the initial image in the image template as a first color system to obtain the target image.
[0099] It is understandable that the first color system mentioned above can be a warm color system, such as red, orange, yellow, etc.
[0100] Optionally, in one embodiment, if the number of different tone values is less than a fifth preset value, the hue type of the image template is determined to be a second color system. The second color system can be a cool color system, such as blue, green, etc.
[0101] Furthermore, in one exemplary embodiment, it can be combined with Figure 5 The above-described image generation apparatus will now be described. Figure 5 This is a structural block diagram of an image generation apparatus according to an embodiment of this application. Figure 2 ).
[0102] like Figure 5 As shown, the image generation device further includes: an acquisition module 52, configured to acquire multiple voice feature dimensions corresponding to the first voice feature information, wherein the multiple voice feature dimensions include at least one of the following: voice volume dimension, voice jitter frequency dimension, and speech rate dimension; acquire multiple psychological stress values of the first voice feature information under the multiple voice feature dimensions; determine abnormal stress values from the multiple psychological stress values; and, if the number of abnormal stress values is greater than a sixth preset value, send a prompt message to the first object to indicate that the first object has the abnormal stress value.
[0103] Among them, the sixth preset value mentioned above can be 0, for example. If the number of abnormal pressure values is determined to be 1, it can be determined that the first object has an abnormal pressure value under a certain feature dimension.
[0104] Furthermore, when the number of abnormal stress values exceeds a sixth preset value, sending a prompt message to the first object may further include: acquiring a first number of the aforementioned multiple psychological stress values and a second number of abnormal stress values; if the ratio of the second number to the first number is determined to be greater than a preset threshold, then a prompt message is sent to the first object to indicate that the first object has the abnormal stress value, and the first object may also be prompted to take stress reduction measures; wherein, the ratio of the second number to the first number indicates the degree of abnormal stress experienced by the first object, and is directly proportional to the degree of stress reduction achieved by the stress reduction measures.
[0105] In other embodiments, the following scheme is also proposed, specifically including: determining the voice volume of the first voice feature information in the voice volume dimension, the voice jitter frequency of the first voice feature information in the voice jitter dimension, and the voice speed of the first voice feature information in the voice speed dimension; and determining a first dimension threshold in the voice volume dimension, a second dimension threshold in the voice jitter dimension, and a third dimension threshold in the voice speed dimension; and determining that the first voice feature information has an abnormal pressure value in the voice speed dimension when the voice volume exceeds the first dimension threshold, or the voice jitter frequency exceeds the second dimension threshold, or the voice speed exceeds the third dimension threshold.
[0106] The method in this embodiment can be used in the most familiar and relaxing home environment for users. It scores and presents the user's mental health through emotion recognition and drawing creation.
[0107] Embodiments of this application also provide a storage medium including a stored program, wherein the program executes any of the methods described above when it is run.
[0108] Optionally, in this embodiment, the storage medium may be configured to store program code for performing the following steps:
[0109] S1, send a preset text to the smart device;
[0110] S2, parse the voice data uploaded by the smart device to obtain the keywords of the voice data and the voice feature information corresponding to the voice data, wherein the voice data is the voice data generated when the first object reads the preset text;
[0111] S3, Generate the target image corresponding to the first object based on the speech feature information and the image template corresponding to the keyword.
[0112] Embodiments of this application also provide an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0113] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0114] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0115] S1, send a preset text to the smart device;
[0116] S2, parse the voice data uploaded by the smart device to obtain the keywords of the voice data and the voice feature information corresponding to the voice data, wherein the voice data is the voice data generated when the first object reads the preset text;
[0117] S3, Generate the target image corresponding to the first object based on the speech feature information and the image template corresponding to the keyword.
[0118] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0119] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.
[0120] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0121] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. An image generation method, characterized in that, include: Send a preset text to a smart device; The voice data uploaded by the smart device is analyzed to obtain the keywords of the voice data and the voice feature information corresponding to the voice data. The voice data is the voice data generated when the first object reads the preset text. Generate a target image corresponding to the first object based on the speech feature information and the image template corresponding to the keyword; The step of generating the target image corresponding to the first object based on the speech feature information and the image template corresponding to the keyword includes: The speech data is subjected to feature recognition to obtain the word category corresponding to the keyword; an image template pre-set for the word category and an image feature adjustment value corresponding to the speech feature information are obtained; a target image is generated based on the image feature adjustment value and the image template pre-set for the word category. Wherein, obtaining the image feature adjustment value corresponding to the speech feature information includes: The voice volume in the voice feature information is obtained, and when the voice volume is greater than a first preset value and less than a second preset value, the first difference between the voice volume and the first preset value is determined as the image feature adjustment value. Based on the image feature adjustment values and the image template pre-set for the word category, a target image is generated, including: The size of the initial image in the image template is adjusted according to the image feature adjustment value to obtain the target image.
2. The image generation method according to claim 1, characterized in that, Obtaining the image feature adjustment value corresponding to the speech feature information further includes: The speech jitter frequency in the speech feature information is obtained, and if the speech jitter frequency is determined to be greater than a third preset value, the second difference between the speech jitter frequency and the third preset value is determined as the image feature adjustment value. Generating a target image based on the image feature adjustment value and an image template pre-set for the word category, further includes: The width of the outline lines of the initial image in the image template is adjusted according to the image feature adjustment value to obtain the target image.
3. The image generation method according to claim 1, characterized in that, Obtaining the image feature adjustment value corresponding to the speech feature information further includes: The speech rate of the speech feature information within a preset time period is obtained, and if the speech rate is determined to be greater than a fourth preset value, the third difference between the speech rate and the fourth preset value is determined as the image feature adjustment value. Generating a target image based on the image feature adjustment value and an image template pre-set for the word category, further includes: The target image is obtained by adjusting the tilt angle of the outline lines of the initial image in the image template according to the image feature adjustment value.
4. The image generation method according to any one of claims 1-3, characterized in that, Generating a target image corresponding to the first object based on the speech feature information and the image template corresponding to the keyword further includes: The number of different pitch values is obtained from the speech feature information; If the number of different tone values is greater than a fifth preset value, the hue type of the initial image in the image template is determined to be the first color system, and the target image is obtained.
5. The image generation method according to any one of claims 1-3, characterized in that, The method further includes: Obtain multiple speech feature dimensions corresponding to the speech feature information, wherein the multiple speech feature dimensions include at least one of the following: speech volume dimension, speech jitter frequency dimension, and speech rate dimension. Obtain multiple psychological stress values of the speech feature information under the multiple speech feature dimensions; Identify abnormal stress values from the plurality of psychological stress values; If the number of abnormal pressure values exceeds a sixth preset value, a prompt message is sent to the first object to indicate that the abnormal pressure value exists.
6. An image generation apparatus, characterized in that, include: The sending module is used to send preset text to smart devices; The parsing module is used to parse the voice data uploaded by the smart device to obtain the keywords of the voice data and the voice feature information corresponding to the voice data. The voice data is the voice data generated when the first object reads the preset text. The determining module is used to generate a target image corresponding to the first object based on the speech feature information and the image template corresponding to the keyword; The determining module is further configured to: The speech data is subjected to feature recognition to obtain the word category corresponding to the keyword; an image template pre-set for the word category and an image feature adjustment value corresponding to the speech feature information are obtained; a target image is generated based on the image feature adjustment value and the image template pre-set for the word category. The determining module is further configured to: The voice volume in the voice feature information is obtained, and when the voice volume is greater than a first preset value and less than a second preset value, the first difference between the voice volume and the first preset value is determined as the image feature adjustment value. Based on the image feature adjustment values and the image template pre-set for the word category, a target image is generated, including: The size of the initial image in the image template is adjusted according to the image feature adjustment value to obtain the target image.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method described in any one of claims 1 to 5.
8. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 5 through the computer program.