Method, system and electronic device for real-time generating dynamic images based on conversation content

By acquiring voice signals to determine interaction parameters, generating and mapping dynamic images to the LED array, the problem of lack of visual feedback in the prior art is solved, and a more immersive human-computer interaction experience is provided.

CN119815135BActive Publication Date: 2025-08-01BEIJING APAILANG CREATIVITY TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411980794.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-08-01
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The existing human-computer dialogue methods lack the generation of rich visual feedback elements in scenes that rely on visual images for content understanding, and cannot provide an immersive user experience.

Method used

By acquiring voice signals to determine interaction parameters, matching visual elements to generate dynamic images, and mapping them to LED arrays, controlling the state changes of LED light emitting units, and designing personalized images in combination with interactive scenes and tasks.

Benefits of technology

It realizes the generation of dynamic images in the process of human-computer interaction, improving the immersive and interactive experience of users, and is especially suitable for interactive experience scenarios such as amusement parks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119815135B_ABST
    Figure CN119815135B_ABST
Patent Text Reader

Abstract

The present application provides a method, a system and an electronic device for generating dynamic images in real time based on conversation content. The method obtains an input voice signal, determines interaction parameters such as an interaction theme, a quantified value of emotional tendency, an interaction time, an interaction location, an interaction role corresponding to the voice signal, and an interaction task preset according to the interaction location. Then, according to the interaction parameters, a combination of visual elements matching each parameter in the interaction parameters is determined, and a moving dynamic image is generated based on each visual object unit designed in advance according to the interaction parameters in the interaction scenario in the combination of visual elements. The dynamic image is mapped onto an LED array for display, and the state change of the corresponding LED lighting unit in the LED array is controlled according to the pixel position of the dynamic image. It can realize a human-computer interaction dialogue around the participants during the interaction process, and generate a matching dynamic image concomitantly, providing a more immersive interaction experience for the participants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, system, and electronic device for real-time generating dynamic images based on conversation content. Background Art

[0002] With the development of artificial intelligence technology, application scenarios of human-computer conversation have become common in all aspects of daily life. Existing human-computer conversation scenarios mainly focus on generating and displaying text content based on human-computer conversation content. For some scenarios that rely on visual images for content understanding, such as cultural and tourism scenarios, museum collection introductions, theme amusement parks, etc., existing human-computer conversation methods can only generate text content and lack the ability to generate richer visual feedback elements, and cannot provide users with a more immersive usage experience and interaction experience. Summary of the Invention

[0003] In view of this, embodiments of this application provide a method, system, and electronic device for real-time generating dynamic images based on conversation content, so as to realize real-time generating dynamic images according to the conversation content of human-computer interaction, thereby providing users with a more immersive usage experience and interaction experience.

[0004] In a first aspect, embodiments of this application provide a method for real-time generating dynamic images based on conversation content, where the method includes:

[0005] Obtain an input voice signal, and determine interaction parameters corresponding to the voice signal, where the interaction parameters include: interaction topic, emotional tendency quantization value, interaction time, interaction location, interaction role, and an interaction task preset according to the interaction location;

[0006] Determine a visual element combination that matches the interaction parameters, and generate a dynamic image based on the visual element combination, where the visual element combination includes multiple visual elements, and the visual elements are visual object units designed in advance according to the interaction parameters in the interaction scenario;

[0007] Map the dynamic image to an LED array, and control the state change of the corresponding LED light-emitting unit in the LED array according to the pixel position of the dynamic image.

[0008] In some possible embodiments, the determining a visual element combination that matches the interaction parameters includes:

[0009] Input the interaction parameters into the visual element matching model, and construct the visual element combination based on each visual element output by the visual element matching model, where the visual element matching model is a data model trained in advance using the matching relationships between interaction theme content tags, sentiment tendency tags, and each visual element as sample data.

[0010] In some possible embodiments, generating the dynamic image based on the visual element combination includes:

[0011] Determine two visually adjacent elements G 01 、G 02 ;

[0012] Based on each pixel point d 01 in the visual element G 01 and the corresponding pixel point d 02 in the visual element G 02 , calculate the deformed pixel point d t within time t through a preset image deformation algorithm;

[0013] Based on the pixel point d 01 , the pixel point d 02 , and the deformed pixel point d t , generate the dynamic image in combination with the playback order.

[0014] In some possible embodiments, the method further includes:

[0015] Perform linear interpolation processing on the pixel point d 01 , the pixel point d 02 , and the deformed pixel point d t using a preset color interpolation algorithm to generate a dynamic color image.

[0016] In some possible embodiments, the method further includes:

[0017] Obtain the initial music that matches the user's identity tag information, play the initial music at the start of the interaction task, and control the state change of the LED lighting unit based on the pixel coordinate values and color values of the respective pixels corresponding to the dynamic image.

[0018] In some possible embodiments, mapping the dynamic image onto the LED array includes:

[0019] Using the nearest neighbor interpolation algorithm, determine the position of the target pixel point in the display image of the LED array corresponding to each pixel point of the dynamic image based on the mapping relationship between the first coordinate range of the dynamic image and the second coordinate range of the display image of the LED array;

[0020] Control the brightness, color, and on / off state of the LED light-emitting unit corresponding to the target pixel point according to the pixel coordinate value and color value corresponding to the position of the target pixel point.

[0021] In some possible embodiments, the method further includes:

[0022] If there is no voice signal input beyond a preset time, control the LED array to enter an energy-saving mode, where the energy-saving mode includes: randomly determining a target visual element for display at a first set display time interval, and controlling the LED array to display the target visual element within a second set display time interval.

[0023] In a second aspect, the present application provides a system for real-time generation of dynamic images based on conversation content, and the system includes:

[0024] An input module, configured to obtain an input voice signal and determine interaction parameters corresponding to the voice signal, where the interaction parameters include: an interaction theme, an emotional tendency quantization value, an interaction time, an interaction location, an interaction role, and an interaction task preset according to the interaction location;

[0025] An image generation module, configured to determine a combination of visual elements that matches the interaction parameters, and generate a dynamic image based on the combination of visual elements, where the combination of visual elements includes multiple visual elements, and the visual elements are visual object units designed in advance according to each interaction parameter in the interaction scenario;

[0026] An image display module, configured to map the dynamic image onto the LED array, and control the state change of the corresponding LED light-emitting unit in the LED array according to the pixel position of the dynamic image.

[0027] In a third aspect, an embodiment of the present application provides an electronic device, where the electronic device includes: a processor; and a memory storing a program; where the program includes instructions that, when executed by the processor, cause the processor to execute the method for real-time generation of dynamic images based on conversation content described in the first aspect.

[0028] In a fourth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium storing computer instructions, characterized in that the computer instructions are used to cause a computer to execute the method for real-time generation of dynamic images based on conversation content described in the first aspect.

[0029] Advantages of the present application:

[0030] The present application provides a method, a system and an electronic device for real-time generating dynamic images based on conversation content. Among them, the method obtains an input voice signal, determines interaction parameters such as an interaction theme, an emotional tendency quantization value, an interaction time, an interaction location, an interaction role corresponding to the voice signal, and an interaction task preset according to the interaction location. Then, according to the interaction parameters, a combination of visual elements matching each parameter in the interaction parameters is determined, and a moving dynamic image is generated based on each visual object unit designed in advance according to the interaction parameters in the interaction scene in the combination of visual elements. Then, the dynamic image is mapped onto an LED array for display, and the state change of the corresponding LED light-emitting unit in the LED array is controlled according to the pixel position of the dynamic image.

[0031] By selecting the embodiments of the present application, the overall experience of playing dynamic images of an LED array for creating a human-machine conversation background can be created. With personalized image design combining the interaction scene and the interaction task, a human-machine interaction conversation around the participants can be realized during the interaction, and a matching dynamic image can be generated concomitantly. For some scenarios that focus on interaction experience such as amusement parks, theme parks, exhibition halls, etc., selecting the embodiments of the present application can provide users with a more immersive usage experience and interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In the following description of exemplary embodiments in conjunction with the drawings, more details, features and advantages of the present application are disclosed. In the drawings:

[0033] Figure 1 FIG. shows a flowchart of a method for real-time generating dynamic images based on conversation content provided by an embodiment of the present application;

[0034] Figure 2 FIG. shows a schematic logical architecture diagram of a system for real-time generating dynamic images based on conversation content provided by an embodiment of the present application;

[0035] Figure 3 FIG. shows another schematic logical architecture diagram of a system for real-time generating dynamic images based on conversation content provided by an embodiment of the present application;

[0036] Figure 4 FIG. shows a structural block diagram of an exemplary electronic device capable of implementing the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] Embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.

[0038] It should be understood that the various steps described in the method embodiments of the present application can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this regard.

[0039] The term "including" and its variations used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0040] It should be noted that the modifications of "one" and "a plurality" mentioned in the present application are illustrative rather than restrictive. Those skilled in the art should understand that, unless clearly stated otherwise in the context, it should be understood as "one or more".

[0041] In order to generate dynamic images adapted to the current interaction scenario in real time according to the content of the human-computer dialogue, so as to provide a more immersive usage experience and interaction experience for the user. The present application provides a method, a system and an electronic device for generating dynamic images in real time based on the dialogue content. Among them, the method for generating dynamic images in real time based on the dialogue content provided by the present application can be applied to any electronic device with the ability to generate real-time dynamic images based on the dialogue content. Similarly, the system for generating dynamic images in real time based on the dialogue content provided by the present application can be deployed in the electronic device with the ability to generate real-time dynamic images based on the dialogue content. The type of the electronic device can be different according to different application scenarios, including but not limited to: personal mobile terminal devices, computers or servers, etc.

[0042] As an example, the electronic device can be a triggering device used to provide interactive services to tourists in a theme park. The triggering device can be a portable interactive device, such as a bracelet or a prop, or a fixed interactive device installed in the amusement park, such as a fixed touch button or a touch screen.

[0043] In a first aspect, the present application provides a method for generating dynamic images in real time based on conversation content. As Figure 1 shown, the method includes the following steps:

[0044] S11. Obtain the input voice signal and determine the interaction parameters corresponding to the voice signal. The interaction parameters include: interaction topic, emotional tendency quantization value, interaction time, interaction location, interaction role, and interaction tasks preset according to the interaction location.

[0045] S12. Determine the visual element combination that matches the interaction parameters, and generate a dynamic image based on the visual element combination. Among them, the visual element combination includes multiple visual elements, and the visual elements are visual object units designed in advance according to the interaction parameters in the interaction scenario.

[0046] S13. Map the dynamic image to the LED array, and control the state change of the corresponding LED lighting unit in the LED array according to the pixel position of the dynamic image.

[0047] In the embodiment of the present application, by obtaining the input voice signal, determining the interaction parameters such as the interaction topic, emotional tendency quantization value, interaction time, interaction location, interaction role, and interaction tasks preset according to the interaction location corresponding to the voice signal, and then determining the visual element combination that matches each parameter in the interaction parameters according to the interaction parameters, and generating a moving dynamic image based on each visual object unit designed in advance according to the interaction parameters in the visual element combination. Then, map the dynamic image to the LED array for display, and control the state change of the corresponding LED lighting unit in the LED array according to the pixel position of the dynamic image.

[0048] By selecting the embodiment of the present application, the overall experience of playing the dynamic image of the LED array in the background of human-computer dialogue can be created. Combining the personalized image design of the interaction scenario and the interaction task, it is possible to achieve a human-computer interaction dialogue around the participant during the interaction process, and generate a matching dynamic image in a companion manner, providing a more immersive interaction experience for the participant.

[0049] To facilitate the understanding of the beneficial effects of the embodiments of the present application, this text takes the method provided by the embodiments of the present application applied to an amusement park as an example. Suppose there are several play attractions in a certain amusement park, and a complete play route is formed among different play attractions. If a tourist arrives at a play attraction on this play route, they can interact with the interactive device in this play attraction to complete the interactive mini-game specially designed for this play attraction. There is a continuity in the plot among the interactive mini-games between different play attractions. This interactive mini-game is the interactive event. The types of interactive events in different play attractions are different, and the interactive events of each play attraction constitute a complete story.

[0050] As the execution entity of the embodiments of the present application, through interacting with the tourists at the play attraction where the interactive device is located, the interactive device can determine how many interactive events the current tourist has completed, as well as the interactive events that need to be completed currently, the theme corresponding to the current play attraction, the current emotional tendency of the tourist, and the tourist's reply based on the type of interactive event, and generate dynamic images that match the current play attraction, the current theme, the current emotional tendency of the tourist, and the type of interactive event, so as to increase the visual interaction experience of the tourist, thereby enhancing the tourist's play experience.

[0051] The following will elaborate on the above steps S11 to S13 in combination with specific examples:

[0052] Among them, when performing step S11, obtaining the input voice signal may be obtaining the voice signal collected by the microphone module inside the device. For example, for the execution entity of the present application, the input voice signal is the voice signal collected by the microphone of the interactive device. As another implementation manner, obtaining the input voice signal may be obtaining the voice signal collected by an external microphone device and sent to the execution entity. Among them, the external sound pickup device and the execution entity of the method of the embodiments of the present application are not the same device, and the two devices are connected by wireless communication. Exemplarily, the two devices are connected by Bluetooth.

[0053] In the embodiments of the present application, the device with the ability to collect voice signals is simply referred to as the sound pickup device. Based on this, in some embodiments, since the application scenario of the method provided by the embodiments of the present application may be a large public place, there will be other sounds in the place in addition to the voices of the dialogue users, that is, there will be noise interference. To ensure the accuracy of the generated sound effects, during the execution of the above step S11, the method further includes:

[0054] S11-1. Obtain the voice signal collected by the target sound pickup device and add identity tag information to the voice signal.

[0055] S11-2. Perform noise reduction processing on the voice signal of the target sound pickup device based on a preset noise reduction algorithm to obtain the voice signal to be parsed after noise reduction.

[0056] Among them, when performing step S11-1, the identity of the interacting user and their interaction role information can be determined by interacting with other identity recognition devices in the application scenario. Exemplarily, the identity information of the user performing the interaction and the associated role information R can be determined by means of near-field communication, such as NFC, RFID, etc., bracelet induction, or face recognition.

[0057] In the embodiment of the present application, the user performing the interaction is the closest to the sound collection device and has the greatest intensity in the stably collected sound. The dialogue voice between the user and the interaction device is obtained in real time. A close-range sound pickup device is used on the hardware, combined with a preset noise reduction algorithm, to ensure that only the voice of one user is recognized as a dialogue with the interaction device at the same time. The preset noise reduction algorithm performs spectral analysis on the collected audio signal, identifies and filters out environmental noise and other irrelevant sound frequency components, thereby improving the accuracy of speech recognition. Among them, the specific type of the preset noise reduction algorithm can be set according to the actual application scenario, and the present application does not make strict limitations.

[0058] Exemplarily, assume that the input voice signal is S voice (t), where t represents time. Through the preset AI noise reduction algorithm, S voice (t) can be converted into the voice signal to be parsed after noise reduction S' voice (t).

[0059] Furthermore, when performing step S11, among the interaction parameters corresponding to the voice signal, the interaction location can be set according to the location where the interaction behavior occurs, which can specifically be the location obtained by GPS positioning within the execution entity or the location of the device with which it interacts. Similarly, to determine the interaction time corresponding to the voice signal, it can be determined according to the system time of the interaction device of the execution entity. Among them, the system time of the interaction device has a regular calibration function to ensure the accuracy of the interaction time.

[0060] When performing step S11, the interaction theme corresponding to the voice signal can be determined according to the location where the interaction device is located. Exemplarily, taking the aforementioned amusement park as the application scenario, different scenarios correspond to different themes. If interaction devices are set in each scenario, the interaction theme corresponding to the interaction device that receives the voice signal can be determined according to the mapping relationship between the interaction device and the scenario.

[0061] As another implementation manner, the interaction theme corresponding to the voice signal can be determined according to the voice content of the interaction voice of the participants. Specifically, the interaction theme of the voice signal can be determined in the following manner:

[0062] S11-3. According to a preset voice-to-text function, convert the input voice signal into target text content;

[0063] S11-4. Based on the target text content, perform semantic analysis to extract the associated words related to the preset interaction theme included in the target text content;

[0064] S11-5. According to the associated words, determine the interaction theme with the highest degree of association with the associated words as the interaction theme corresponding to the voice signal.

[0065] Among them, in step S11-3, the preset voice-to-text function can be any algorithm function with voice-to-text capabilities. Exemplarily, a trained neural network model for voice-to-text can be used to convert the input voice signal into target text content, or convert the input voice signal to be parsed into target text content. Assume that the preset voice-to-text function is T trans () Then the voice signal S' to be parsed voice (t) can be converted into the corresponding target text content W, where W = T trans (S' voice (t)).

[0066] Furthermore, perform step S11-4. As an implementation manner, through a preset natural language processing tool or model, perform semantic analysis on the target text content W, and extract each noun and the mood-related auxiliary words included in the target text content W, so as to subsequently determine the corresponding interaction theme and emotional tendency quantization value based on the extracted nouns and mood-related auxiliary words. As another implementation manner, based on the applicable scenario of the method provided in the embodiments of the present application, the possible interaction themes can be preset in advance, and such preset interaction themes are the preset interaction themes. Exemplarily, taking the aforementioned amusement park as an example, assume that the applied amusement park is a technology amusement park, which is subdivided into aerospace technology, biotechnology, computer technology, quantum technology, etc., and then match the extracted nouns with the set vocabulary, extract the associated associated words included in the target text content, and further determine which technology theme it is currently in.

[0067] As another implementation, a language text recognition and classification model can be pre-trained based on specific application scenarios, such as a text classification model based on word vectors and deep learning, to determine the interaction theme with the highest degree of relevance to the conversation content. Then, step S11-5 is executed to determine the interaction theme with the highest degree of association with the associated vocabulary as the interaction theme corresponding to the voice signal according to different associated vocabulary. Exemplarily, taking the aforementioned amusement park as the application scenario, assuming that the amusement park is divided into different plot areas, and each plot area corresponds to a theme type, it is possible to determine which plot-related vocabulary is specifically included in the current target text content, and determine the plot with the highest degree of association as the theme type corresponding to the current voice signal.

[0068] Based on the large target text content converted from the participant's voice signal, by constructing a theme word library corresponding to different themes with the vocabulary that different themes may involve, and then analyzing the conversation theme C of the current interaction with the help of a semantic analysis model. Specifically, assuming the text content is C and the theme word library is B, for each word c in C i , calculate the semantic similarity s(c i , b j ) between each word c i and each theme word b j in B. Specifically, the cosine similarity between c i and b j can be calculated based on the word vector model. Then, count the sum of the occurrence frequency f(c i ) and the semantic association degree of each theme-related word, that is, calculate the sum S j of the occurrence frequency and the semantic association degree according to the following formula:

[0069]

[0070] Then, select the theme with the largest S j as the interaction theme corresponding to the voice signal.

[0071] As another implementation, a language text recognition and classification model can be pre-trained based on specific application scenarios, such as a text classification model based on word vectors and deep learning, to determine the interaction theme with the highest degree of relevance to the conversation content. Then, according to different associated vocabulary, determine the interaction theme with the degree of association with the associated vocabulary as the interaction theme corresponding to the voice signal. The specific algorithm implementation in the above steps S11-4 and S11-5 can be understood as: assuming that a theme classification function trained based on the actual application scenario is: F topic (), at this time, input the target text content into this F topic () function, and the corresponding interaction theme can be obtained, that is, C = F topic(W). Among them, the specific interaction topics can be further divided into aerospace technology, biotechnology, computer technology, quantum technology, and so on.

[0072] In some possible embodiments, the quantified value of the emotional tendency corresponding to the voice signal can also be obtained by analyzing the interactive voice signal. As an implementation manner, semantic analysis can be performed based on the target text content, and then the modal particles included in the target text content can be extracted to further assist in judging the current mood and emotional tendency of the participant. Specifically, the quantified value of the emotional tendency can be obtained through the following manner:

[0073] S11-6. Obtain the audio features of the voice signal, where the audio features include: speech rate v, volume l, and pitch p;

[0074] S11-7. Based on the reference speech rate, reference volume, and reference pitch, perform normalization processing on the voice features to obtain a normalized audio difference ratio;

[0075] S11-8. According to the audio difference ratio and in combination with the preset emotional tendency label information, calculate the target emotional tendency score, and determine the score based on the target emotional tendency score as the quantified value of the emotional tendency corresponding to the voice signal.

[0076] Among them, when performing step S11-6, as an implementation manner, the speech rate v of the voice signal can be calculated by calculating the number of words or syllables per unit time in the voice signal. Specifically, by identifying the word or syllable boundaries in the voice signal, then measuring the time interval between the word or syllable boundaries, and then calculating the number of words or syllables per unit time.

[0077] As an implementation manner, the amplitude of the voice signal can be extracted by performing audio feature extraction on the voice signal, and the amplitude of the voice signal is the volume l. Specifically, the corresponding volume can be calculated by calculating the sum of the amplitudes of all sampling points in a sound frame, or the volume of the voice signal can be obtained by calculating the sum of the squares of the amplitudes of all sampling points and then taking the logarithm.

[0078] As an implementation manner, the fundamental frequency of the voice signal can be extracted by performing spectrum analysis on the voice signal through methods such as short-time Fourier transform, and then the fundamental frequency is determined as the pitch p of the voice signal.

[0079] In step S11-7, the reference speech rate v0, reference volume l0, and reference pitch p0 are the standards for measuring the speech rate, volume, and pitch of a speech signal. By comparing the speech rate v of the speech signal with the reference speech rate v0, the speaking speed of the participant can be determined. Similarly, by comparing with the reference volume, the volume level of the participant's speech can be determined, and by comparing with the reference pitch, the pitch level of the participant's speech can be determined. Since the speech rate, volume, and pitch are data under different dimensions, for the convenience of unified calculation, in the embodiments of the present application, the actual speech characteristics of the participant are normalized based on the reference quantity to obtain a normalized audio difference ratio. Among them, corresponding to the audio characteristics of the speech, the audio difference ratio also includes three categories: speech rate audio difference ratio, volume audio difference ratio, and pitch audio difference ratio.

[0080] In the embodiments of the present application, the normalization process can obtain the ratio of the difference between the actual audio characteristic and the reference value to the reference value as the audio difference ratio. Specifically, taking the volume as an example, the volume audio difference ratio r l =(actual volume l of the speech signal - reference volume l0)÷reference volume l0. By analogy, the speech rate audio difference ratio r v =(actual speech rate v of the speech signal - reference speech rate v0)÷reference speech rate v0, and the pitch audio difference ratio r p =(actual pitch p of the speech signal - reference pitch p0)÷reference pitch p0.

[0081] Further, by executing step S11-8, according to the normalized audio difference ratio and in combination with the mapping relationship between the audio difference ratio and the preset emotional tendency label, the target emotional tendency score can be calculated. As an implementation manner, the volume audio difference ratio, the speech rate audio difference ratio, and the pitch audio difference ratio can be multiplied to obtain a comprehensive audio difference weighting coefficient K, that is, K=(1 + volume audio difference ratio r l )×(1 + speech rate audio difference ratio r v )×(1 + pitch audio difference ratio r p ).

[0082] Among them, the preset sentiment tendency label is the data label E that is set in advance step by step according to the sentiment tendency from negative to positive. Among them, the levels corresponding to the specific data labels can be the three simple levels of negative, neutral, and positive, or the nine levels of pessimistic, relatively pessimistic, negative, relatively negative, neutral, relatively positive, positive, relatively optimistic, and optimistic. Different levels correspond to different sentiment tendency label values, that is, different E values. In this way, the target sentiment tendency score E' can be calculated according to the formula E' = K × E. When the participant first uses the service for generating real-time dynamic images based on the conversation content provided by the embodiments of the present application, the sentiment tendency score of the participant can be calculated by performing the above steps S11-6 to S11-8.

[0083] As another implementation manner, with the multiple human-computer interactions of the participant, the speaking speeds v of the participant at different times can be compared i , volume l i and pitch p i to learn the voice characteristics of the participant and gradually perform normalization processing, which can make the calculation of the above comprehensive audio difference ratio K more conform to the conventional voice characteristics of the participant, avoid the comprehensive audio difference ratio obtained by single calculation from overly affecting the result of sentiment tendency analysis, and finally update the sentiment tendency score E' in the above manner.

[0084] Specifically, statistical methods such as mean and standard deviation can be used to update the values of volume, speech rate, and pitch for determining sentiment tendency. Exemplarily, the mean value of volume can be calculated by calculating the mean and standard deviation speech rate mean pitch mean volume standard deviation σ l speech rate standard deviation σ v pitch standard deviation σ p , and then with the help of the mean value of volume speech rate mean pitch mean volume standard deviation σ l speech rate standard deviation σ v pitch standard deviation σ p to perform normalization processing again to obtain the normalized volume value r l ′, the normalized speech rate value r v ′, the normalized pitch value r p ′, and then with the help of the normalized volume value r' l , the normalized speech rate value r' v , and the normalized pitch value r' p calculate the final comprehensive K', making the result of K' more reasonable. Among them, the normalized volume value r' l , the normalized speech rate value r'v , the normalized pitch value r′ p The calculation normalization methods are the same. Here, the speech rate is taken as an example for illustration, and the other two items will not be elaborated:

[0085] Assume that the speech rates of multiple interactive voices of the participant are {v1, v2,..., v n}}, and then calculate their average speech rate It can be calculated according to the following formula

[0086]

[0087] where n is the total number of interactive voices participating in the calculation of the average speech rate, i is the serial number of the interactive voice participating in the calculation of the average speech rate, and then the following formula is used for normalization processing to obtain the normalized speech rate value r v ′:

[0088]

[0089] Then calculate K′ through the following formula:

[0090] K′ = (1 + r l ′)(1 + r v ′)(1 + r p ′)

[0091] In this way, it can be avoided that the conventional speech features of the participant overly affect the calculation accuracy of the emotional tendency quantization value.

[0092] In the embodiment of the present application, when determining the interactive role corresponding to the voice signal in step S11, it can be obtained by parsing the reply voice signal obtained by asking the participant about the current interactive role of the participant during the human-computer interaction process. As another implementation manner, it can be determined by obtaining the interactive role checked by the participant in the interactive interface when the participant enters the interactive scene. Exemplarily, taking the application scenario as an amusement park, when a tourist enters the amusement park, the tourist can check the role they want to play, and the role checked by the tourist at this time is the interactive role. In the embodiment of the present application, when determining the interactive task corresponding to the voice signal in step S11, the corresponding interactive task can be determined by obtaining the position of the execution main body interactive device.

[0093] Based on the multiple interactive parameters obtained in step S11, step S12 can be further executed to determine the visual element combination that matches the interactive parameters. Specifically, as a possible implementation manner, step S12 can be implemented through the following steps:

[0094] S12-1. Input the interaction parameters into the visual element matching model, and construct the visual element combination based on each visual element output by the visual element matching model. The visual element matching model is a data model pre-trained with the matching relationships between interaction theme content tags, sentiment tendency tags, and each visual element as sample data.

[0095] Specifically, each visual element can be designed and produced according to the overall environment of the interaction content and interaction scenario. The visual element can be a geometric figure, a cartoon image of an interaction character, or an IP (intellectual property) figure of a person, an item, etc. It may also be a figure derived from different color combinations of the above figures. With the visual elements in different interaction scenarios, interaction theme content tags are constructed for each visual element in advance according to different interaction scenarios and interaction themes, and the interaction scenarios applicable to the visual element are marked through the interaction theme content tags. Similarly, sentiment tendency tags are constructed for each visual element in advance according to different emotional tendencies, and the emotional state applicable to the visual element is marked through the sentiment tendency tags.

[0096] Then, the matching relationships between each visual element and the interaction theme and interaction sentiment tendency can be reflected by the interaction theme content tags and sentiment tendency tags. Then, using the matching relationships between the interaction theme content tags, sentiment tendency tags, and each visual element as training sample data, the pre-set learning model is trained. Through this learning model, the correlation between the visual element and the applied interaction scenario and the adapted emotional state is learned until the learning model can accurately match the adapted visual element according to the input interaction theme and sentiment tendency. At this time, the trained learning model is the visual element matching model.

[0097] Specifically, during the training process, assume the training data is: {(C1, E1, G1), (C2, E2, G2),..., (C m , E m , G m )}, where C i corresponds to the interaction theme content tag, E i corresponds to the sentiment tendency tag, and G i corresponds to the visual element matching the interaction theme content tag C i and the sentiment tendency tag E i . The model is trained with this training data to optimize the weights of the model, so that the trained model can accurately output the corresponding G based on the input C and E.

[0098] Among them, information such as the position L of the interaction device, the interaction role R, the interaction time t, and the interaction task O can be used as input parameters of the model and input into the model, and the model retrieves the matching visual elements. Among them, the position L can be a coordinate vector (x, y, z) or a region identifier, the interaction role R can be the number of the role type, the interaction time t can be a timestamp or a time period identifier, and the interaction task O can be information such as the number of the preset basic event type, etc.

[0099] Further, a dynamic image is generated based on the obtained visual element combination through the following steps:

[0100] S12-2. Determine two adjacent visual elements G 01 , G 02 ;

[0101] S12-3. Based on each pixel point d 01 in the visual element G 01 and each pixel point d 02 corresponding to the visual element G 02 , calculate the deformed pixel point d t within the time t through a preset image deformation algorithm;

[0102] S12-4. Based on the pixel point d 01 , the pixel point d 02 and the deformed pixel point d t , generate the dynamic image in combination with the playback order.

[0103] As an implementation manner, assume that the above visual element matching model is M, then the output visual element of the visual element matching model M is G = M(L, R, t, O), that is, the visual element output by the matching model M is from a set of visual elements {G 01 , G 02 ,..., G M}.

[0104] Further, the visual elements G 01 , G 02 that need to be displayed adjacent to each other can be determined from the visual element combination, and then any pixel point d 01 of G 01 to the corresponding pixel point d 02 of G 02 is determined, and through a preset image deformation algorithm, the pixel point d t of the two pixel points at the time point t during the deformation process is calculated.The position where the value of t can be flexibly set according to actual requirements. As a preferred implementation, the value of t can be 0 ≤ t ≤ 1s. The preset image deformation algorithm can be flexibly selected from existing image deformation algorithms according to the requirements of the actual application scenario, and this application does not strictly limit the preset image deformation algorithm.

[0105] Specifically, assume that the coordinates of the pixel point in G 01 are d 01 (x 01 , y 01 ), and the coordinates of the corresponding pixel point in G 02 are d 02 (x 02 , y 02 ). With the help of the image deformation algorithm, calculate the coordinates x t and y t of the deformed pixel point d t according to the following formula:

[0106]

[0107] Furthermore, the change order of the entire dynamic image can be understood as starting from the pixel point d 01 in the visual element G 01 , using the deformed pixel point d t as the transition pixel point, and changing to the pixel point d 02 in the visual element G 02 . In this way, a dynamic image d 01 —>d t —>d 02 can be generated based on the above playback order. As another implementation, the visual element frames constructed by each deformed pixel point can be used as key frames, and the deformed pixel point can be inserted between the visual element G 01 and the visual element G 02 through key frame interpolation technology to achieve a dynamic effect.

[0108] In the embodiment of this application, during the execution of the above step S12, a dynamic image can also be generated through the following steps:

[0109] Use a preset color interpolation algorithm to perform linear interpolation processing on the pixel point d 01 , the pixel point d 02 and the deformed pixel point d t , calculate the coordinate values and RGB color values of the corresponding pixel points, and generate a dynamic image.

[0110] In some possible embodiments, during the process of playing the dynamic image generated in step S12 and controlling the state change of the LED lighting unit in step S13, the dynamic image can also be matched with appropriate music in the following manner:

[0111] Obtain the initial music that matches the user's identity tag information. At the start of the interaction task, play the initial music, and control the state change of the LED lighting unit based on the pixel coordinate values and color values of each pixel point corresponding to the dynamic image.

[0112] In the embodiments of the present application, the initial music that matches the user's identity tag can be obtained by matching based on a pre-trained music matching model, or when the user selects an interaction task, different interaction tasks correspond to different background music, and this background music is the initial music. Then, during the process of the user completing the interaction task, by interacting with the user, obtain the user's identity tag, and then determine the interaction task and interaction role selected by the user through the user's identity tag, and then determine the appropriate background music as the initial music based on the interaction task and interaction role. In this way, when the interaction task starts to be executed, while playing the initial music, control the state change of the LED lighting units in the LED array based on the pixel coordinate values, color values, etc. of each pixel point corresponding to the dynamic image, so as to provide the participants with a dual auditory and visual participation experience.

[0113] As another implementation method, during the entire interaction process of the participant, the interactive voice signal can be sampled and analyzed at a set sampling interval. As a preferred implementation method, the set sampling time interval is △t = 3s, and the text content corresponding to the voice signal sampled at time t n is C n , and based on this text content, analyze the interactive theme keyword B n and the emotional tendency quantization value E n . Then, based on the analysis result, generate the visual element G n = M(B n , E n ). Thus, it is possible to form an effect of dynamically switching between the visual elements of the previous time interval and the visual elements of the next time interval, which can help create a more rich immersive and personalized participation experience.

[0114] In some possible embodiments, when performing the above step S13, the dynamic image can be mapped onto the LED array in the following manner:

[0115] S13-1. Use the nearest neighbor interpolation algorithm to determine the positions of the target pixels corresponding to the pixel points of the dynamic image in the display image of the LED array based on the mapping relationship between the first coordinate range of the dynamic image and the second coordinate range of the display image of the LED array;

[0116] S13-2. Control the brightness, color, and on / off state of the LED light-emitting unit corresponding to the target pixel according to the pixel coordinate value and color value corresponding to the target pixel position.

[0117] During the execution of the above steps S13-1 and S13-2, it can be understood that the generated dynamic image is mapped onto LED arrays with different resolutions using the nearest neighbor interpolation algorithm. Specifically, assume that the coordinate range of the image G of the original dynamic image is (0, 0) to (W, H), where W is the width of the original image and H is the height of the original image. The display image of the LED array is the target image to which it needs to be mapped, and the coordinate range of this target image G d is (0, 0) to (W d , H d ). W d is the width of the target image, and H d is the height of the target image. Further, for the original image and the target image, there is a corresponding mapping relationship between the pixel points. Specifically, for the pixel point coordinates (x, y) in the original image, the corresponding pixel point coordinates (x d , y d ) in the target image satisfy the following formula:

[0118]

[0119] Assume that the RGB color value of the pixel point d 01 in G 01 is (r 01 , g 01 , b 01 ), and the RGB color value of the pixel point d 02 in G 02 is (r 02 , g 02 , b 02 ). The calculation of the RGB color value of the deformed pixel point d t at time t satisfies the following formula:

[0120]

[0121] Further, the coordinate value of the actual LED light-emitting unit corresponding to the pixel point (x d , y d , y d ) in the target image G is determined by xd 、y d The color value of the light emitting unit is obtained by rounding down. d 、g d 、b d The rounding operation is performed.

[0122] Thus, during the execution of step S13-2, the brightness, color, and on / off state of the LED light-emitting unit can be controlled according to the calculated pixel coordinate values and color values, thereby achieving the effect of generating and displaying a matching dynamic image according to the interactive voice signal. In the embodiment of the present application, the LED light-emitting unit includes: LED lamp beads, organic light-emitting diodes (OLEDs), and light-emitting units obtained by various LED chip packaging methods.

[0123] In some possible embodiments, the interactive voice signal can be sampled at shorter and higher frequency sampling intervals, and then corresponding visual elements can be generated, thereby forming a closer dynamic image matching effect with the interpersonal dialogue. In this process, assuming that the sampling function is F sample (t), at t n The text content sampled at the sampling time is C n =F sample (t n ), then based on the text content C n Matching generates visual elements G n If a sampling error occurs during a certain interactive voice conversation, the visual elements and dynamic display effects from the previous sampling interval are retained and played back in a loop until the next sampling interval, when the speech signal is successfully recognized and the corresponding visual elements and dynamic display effects are successfully generated. This reduces the negative impact of sampling failures on the visual experience.

[0124] In some possible embodiments, the voice signal input can be monitored. If there is no voice signal input for a predetermined time, the LED array is controlled to enter the energy-saving mode. Specifically, the energy-saving mode includes: a ), every t a Time from the collection of visual elements {G 01 ,G 02 ,...,G M} randomly selects visual elements to display. Then controls the LED array to display the second set time length (assuming t b ), at t bDisplay the selected visual elements within a period of time. Among them, the specific time lengths of the first set time and the second set time can be flexibly set according to the actual scenario. Exemplarily, assuming no interaction occurs, visual elements in the visual element library are randomly selected every 10 seconds and displayed on the LED array, and the specific display duration is 3 seconds. In this way, it can not only play an energy-saving role but also attract participants to participate in the interaction.

[0125] In a second aspect, the present application provides a system for real-time generating dynamic images based on conversation content. Among them, as Figure 2 shown, the system 20 includes:

[0126] An input module 201, configured to obtain an input voice signal and determine interaction parameters corresponding to the voice signal, where the interaction parameters include: an interaction topic, an emotional tendency quantization value, an interaction time, an interaction location, an interaction role, and an interaction task preset according to the interaction location;

[0127] An image generation module 202, configured to determine a combination of visual elements matching the interaction parameters and generate a dynamic image based on the combination of visual elements. Among them, the combination of visual elements includes multiple visual elements, and the visual elements are visual object units designed in advance according to various interaction parameters in the interaction scenario;

[0128] An image display module 203, configured to map the dynamic image to the LED array and control the state change of the corresponding LED light-emitting unit in the LED array according to the pixel position of the dynamic image.

[0129] In some possible implementation manners, the system for real-time generating dynamic images based on conversation content provided by the present application may also be as Figure 3 shown, and specifically implement the method described in the first aspect through the following several modules:

[0130] A participant identity information recognition module 31, configured to perform the above steps S11-1 and S11-2, identify the participant's identity, and add identity tag information to the voice signal, so as to match different visual elements according to different participants subsequently.

[0131] A conversation content analysis module 32, configured to perform the above steps S11-3 to S11-5 to determine the text content corresponding to the voice signal and the interaction topic included in the text content.

[0132] An emotional tendency analysis module 33, configured to perform the above steps S11-6 to S11-8 to determine the emotional tendency quantization value corresponding to the voice signal.

[0133] The visual element and dynamic image generation module 34 is used to execute the above steps S12-1 to S12-4 to determine the matching visual elements and generate dynamic images based on the matched visual elements.

[0134] The LED array control and display output module 35 is used to execute the above step S13 to control the display of the LED array based on the generated dynamic images and the interaction duration.

[0135] By selecting the embodiments of the present application, it is possible to generate dynamic images that match comprehensive information such as human-machine dialogue, the location of the interaction device, interaction roles, interaction time, and interaction tasks in an accompanying and real-time manner, and create an immersive and interactive participation experience.

[0136] Among them, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the present application comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0137] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only for illustrative purposes and do not limit the scope of these messages or information.

[0138] In a third aspect, an exemplary embodiment of the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and the computer program, when executed by the at least one processor, causes the electronic device to execute the method according to the embodiments of the present application.

[0139] An exemplary embodiment of the present application further provides a non-transitory computer-readable storage medium storing a computer program, where the computer program, when executed by a processor of a computer, causes the computer to execute the method according to the embodiments of the present application.

[0140] An exemplary embodiment of the present application further provides a computer program product, including a computer program, where the computer program, when executed by a processor of a computer, causes the computer to execute the method according to the embodiments of the present application.

[0141] Reference Figure 4, the structural block diagram of the electronic device 400 that can be used as the server or client of the present application will now be described. It is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0142] As Figure 4 shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM 402) or the computer program loaded from the storage unit 408 into the random access memory (RAM 403). In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. The input / output interface (I / O interface 405) is also connected to the bus 404.

[0143] Multiple components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, an output unit 407, a storage unit 408, and a communication unit 409. The input unit 406 can be any type of device that can input information into the electronic device 400. The input unit 406 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 407 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 408 can include but is not limited to a magnetic disk, an optical disk. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0144] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above. For example, in some embodiments, the aforementioned method of real-time generating dynamic images based on conversation content can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. In some embodiments, the computing unit 401 can be configured to execute the aforementioned method of real-time generating dynamic images based on conversation content by any other suitable means (e.g., by means of firmware).

[0145] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0146] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0147] As used in this application, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0148] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0149] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a front-end component (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0150] A computer system can include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other.

Claims

1. A method for real-time generating dynamic images based on conversation content, characterized in that, An interaction device applied to an interaction scenario, the interaction scenario including: a playground, the method including: Obtain an input voice signal, and determine interaction parameters corresponding to the voice signal, the interaction parameters including: an interaction theme, an emotional tendency quantization value, an interaction time, an interaction location, an interaction role, and an interaction task preset according to the interaction location, the interaction location being different locations in the interaction scenario, and the interaction task being a preset basic event type; Determine a visual element combination matching the interaction parameters, and generate a dynamic image based on the visual element combination, wherein the visual element combination includes a plurality of visual elements, and the visual elements are visual object units designed in advance according to each interaction parameter in the interaction scenario; Map the dynamic image onto an LED array, and control the state change of the corresponding LED light-emitting unit in the LED array according to the pixel position of the dynamic image, including: Adopt a nearest neighbor interpolation algorithm, and determine the position of the target pixel corresponding to each pixel point of the dynamic image in the display image of the LED array based on the mapping relationship between the first coordinate range of the dynamic image and the second coordinate range of the display image of the LED array; Control the brightness, color, and on / off of the LED light-emitting unit corresponding to the target pixel according to the coordinate value and color value corresponding to the target pixel position.

2. The method according to claim 1, wherein The determining of the visual element combination matching the interaction parameters includes: Input the interaction parameters into a visual element matching model, and construct the visual element combination based on each visual element output by the visual element matching model, wherein the visual element matching model is a data model trained in advance using the matching relationship between the interaction theme content label, the emotional tendency label, and each visual element as sample data.

3. The method according to claim 1, wherein The generating of the dynamic image based on the visual element combination includes: Determine two visually adjacent displayed visual elements G in the combination of the visual elements 01 , G 02 ; Based on the visual element G 01 for each pixel point d 01 in the visual element G 02 corresponding to each pixel point d 02 , calculate the deformed pixel point d within time t through a preset image deformation algorithm t ; Based on the pixel point d 01 and the pixel point d 02 as well as the deformed pixel point d t to generate the moving image in combination with the playback order.

4. The method according to claim 3, wherein The method further includes: Perform linear interpolation processing on the pixel point d 01 , the pixel point d 02 and the deformed pixel point d t to generate a dynamic color image.

5. The method according to claim 3, wherein The method further includes: Obtain an initial music matching the user's identity label information, play the initial music when the interaction task starts, and control the state change of the LED light-emitting unit based on the pixel coordinate value and color value of each pixel point corresponding to the dynamic image.

6. The method according to claim 1, wherein The method further includes: If there is no voice signal input beyond a preset time, control the LED array to enter an energy-saving mode, wherein the energy-saving mode includes: randomly determining a target visual element for display at a first set display time interval, and controlling the LED array to display the target visual element within a second set display time interval.

7. A system for real-time generation of dynamic images based on conversation content, characterized in that, An interaction device applied to an interaction scenario, the interaction scenario including: a playground, the system including: An input module, configured to obtain an input voice signal, and determine interaction parameters corresponding to the voice signal, the interaction parameters including: an interaction theme, an emotional tendency quantization value, an interaction time, an interaction location, an interaction role, and an interaction task preset according to the interaction location, the interaction location being different locations in the interaction scenario, and the interaction task being a preset basic event type; An image generation module, configured to determine a combination of visual elements that matches the interaction parameters, and generate a dynamic image based on the combination of visual elements, wherein the combination of visual elements includes a plurality of visual elements, and the visual elements are visual object units designed in advance according to the interaction parameters in the interaction scenario; An image display module, configured to map the dynamic image onto an LED array, and control the state change of the corresponding LED light-emitting units in the LED array according to the pixel positions of the dynamic image; Specifically, the image display module is configured to: determine the positions of the target pixel points corresponding to the pixel points of the dynamic image in the display image of the LED array based on the mapping relationship between the first coordinate range of the dynamic image and the second coordinate range of the display image of the LED array by using the nearest neighbor interpolation algorithm; Control the brightness, color, and on / off state of the LED light-emitting unit corresponding to the target pixel point according to the coordinate values and color values corresponding to the target pixel point position.

8. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing a program; wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method according to any one of claims 1-6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Human-computer interaction method, device and equipment and storage medium

    CN117111738A