System, method and device for generating LED dot matrix image

By introducing voice interaction and multimodal AI models into the LED dot matrix screen system, voice data is converted into image data, solving the problems of single interaction mode and low generation efficiency of the existing LED dot matrix screen, and intelligent LED dot matrix image generation and dynamic display are realized.

CN120223945APending Publication Date: 2025-06-27SHANGHAI FORTUNE TECHGROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510497012.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing LED dot matrix screens need to rely on APP, remote control or fixed programs to adjust LED dot matrix images, resulting in a single interaction method and a cumbersome process of generating LED dot matrix images, which makes it impossible to achieve real-time dynamic generation of personalized images.

Method used

A system for generating LED dot matrix images is proposed, including a control device and a cloud server. Voice data is obtained through a voice acquisition device, and voice data is converted into image data using a multimodal AI model to realize automatic LED dot matrix image generation and dynamic display based on voice interaction.

Benefits of technology

The intelligentization of LED dot matrix screen is realized, convenient voice interaction methods are provided, and the generation efficiency of LED dot matrix images is improved, and the problems of single interaction methods and low generation efficiency in the prior art are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223945A_ABST
    Figure CN120223945A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses a system, method and device for generating an LED dot matrix image. Sending first voice data acquired by the voice acquisition device to a cloud server based on a first transmission protocol through the control device; the cloud server calls a first conversion model to convert the first voice data into first text data and returns the first text data to the control device; the control device sends the first text data to a cloud server based on a second transmission protocol; the cloud server calls a second conversion model and a third conversion model in sequence, converts the first text data into a first text graph description text, generates first image data and returns the first image data to the control device; the control device processes the first image data to obtain processed first image data suitable for being displayed by the LED dot matrix screen and sends the processed first image data to the LED dot matrix screen to be displayed; the generation efficiency of the LED dot matrix image can be improved, and a more convenient way of interaction with the LED dot matrix screen is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of image processing, and in particular, to a system, method and device for generating an LED dot matrix image. Background Art

[0002] A light-emitting diode (LED) dot matrix screen is a display device composed of a large number of LEDs arranged in a matrix form. By controlling the on / off state of each LED, text, patterns, animations or video content can be displayed. At this time, the image content displayed on the LED dot matrix screen is the LED dot matrix image, that is, the LED dot matrix image is a static or dynamic visual graphic formed by controlling the on / off of LEDs at different positions on the LED dot matrix screen.

[0003] Traditional methods for generating LED dot matrix images include the following:

[0004] 1. At least one image to be displayed is pre-stored in the LED dot matrix device. The control unit in the LED dot matrix device selects one of the images to be displayed through an instruction, and controls the LED dot matrix screen in the LED dot matrix device to display, so as to generate an LED dot matrix image.

[0005] 2. A communication connection (such as a Bluetooth connection or a WiFi connection) is established between the LED dot matrix device and the user terminal. Based on this communication connection, the image to be displayed in the user terminal is sent to the LED dot matrix device, so as to be displayed through the LED dot matrix screen in the LED dot matrix device to generate an LED dot matrix image.

[0006] 3. The LED dot matrix device is connected to sensors (such as an image sensor, a temperature sensor, etc.). The control unit in the LED dot matrix device controls the LED dot matrix screen in the LED dot matrix device to switch the currently displayed LED dot matrix image based on the input data collected by the sensors.

[0007] However, existing LED dot matrix screens need to rely on an APP, a remote control or a fixed program to adjust the LED dot matrix image, resulting in a single interaction method and a cumbersome process for generating the LED dot matrix image. At the same time, the LED dot matrix screen cannot generate personalized images according to the real-time needs of users. When users need to change the LED dot matrix image, they often need to operate manually and cannot achieve real-time dynamic generation of the LED dot matrix image. Summary of the Invention

[0008] In view of this, the present disclosure proposes a system, method and device for generating an LED dot matrix image, which can improve the intelligent level of the LED dot matrix screen and realize automatic generation and dynamic display of the LED dot matrix image based on voice interaction.

[0009] According to one aspect of the present disclosure, a system for generating an LED dot matrix image is provided. The system includes a control device and a cloud server;

[0010] The control device is communicatively connected to the LED dot matrix screen, and the control device is connected to a voice collection device. The control device is configured to: obtain first voice data collected by the voice collection device, where the first voice data is used to describe the LED dot matrix image to be displayed on the LED dot matrix screen; send the first voice data to the cloud server that has established a communication connection with the control device based on a first transmission protocol;

[0011] The cloud server is configured to: when receiving the first voice data based on the first transmission protocol, call a first conversion model to convert the first voice data into first text data; and send the first text data to the control device based on the first transmission protocol;

[0012] The control device is further configured to: when receiving the first text data based on the first transmission protocol, send the first text data to the cloud server based on a second transmission protocol;

[0013] The cloud server is further configured to: when receiving the first text data based on the second transmission protocol, sequentially call a second conversion model and a third conversion model to convert the first text data into a first text-to-image description text and then generate first image data; send the first image data to the control device based on the second transmission protocol;

[0014] The control device is further configured to: when receiving the first image data based on the second transmission protocol, process the first image data to obtain processed first image data suitable for display on the LED dot matrix screen; send the processed first image data to the LED dot matrix screen for display.

[0015] In a possible implementation, the control device is further configured to: after sending the processed image data to the LED dot matrix screen for display, if second voice data is received, send the second voice data to the cloud server based on the first transmission protocol; where the second voice data is used to describe the content to be modified in the LED dot matrix image currently displayed on the LED dot matrix screen;

[0016] The cloud server is further configured to: when receiving the second voice data based on the first transmission protocol, call a first conversion model to convert the second voice data into second text data; and send the second text data to the control device based on the first transmission protocol;

[0017] The control device is further configured to: when receiving the second text data based on the first transmission protocol, send the second text data to the cloud server based on the second transmission protocol;

[0018] The cloud server is further configured to: when receiving the second text data based on the second transmission protocol, sequentially call the second conversion model and the third conversion model to convert the second text data into a second text-to-image description text and then generate second image data; send the second image data to the control device based on the second transmission protocol;

[0019] The control device is further configured to: when receiving the second image data based on the second transmission protocol, process the second image data; send the processed second image data to the LED dot matrix screen to modify the LED dot matrix image currently displayed on the LED dot matrix screen.

[0020] In a possible implementation manner, the cloud server is further configured to: when the second voice data instructs to modify a local area in the LED dot matrix image currently displayed on the LED dot matrix screen, after sequentially calling the second conversion model and the third conversion model to generate the second image data, obtain the local pixel positions of the local area in the second image data, and send the local pixel positions to the control device;

[0021] The control device is further configured to: map the image content at the local pixel positions in the second image data to the corresponding screen positions of the local area on the LED dot matrix screen to obtain the processed second image data; send the processed second image data to the LED dot matrix screen to modify the LED dot matrix image currently displayed on the LED dot matrix screen.

[0022] In a possible implementation manner, the control device is further configured to:

[0023] After sending the processed image data to the LED dot matrix screen for display, output an image modification prompt, where the image modification prompt is used to prompt the user to determine whether to modify the LED dot matrix image currently displayed on the LED dot matrix screen, and output the second voice data when the LED dot matrix image needs to be modified.

[0024] In a possible implementation manner, the control device is further configured to:

[0025] Determine whether the audio data collected by the voice acquisition device includes a preset keyword to determine whether the audio data is target voice data, where the target voice data includes the first voice data and the second voice data;

[0026] When the audio data includes the preset keyword corresponding to the first voice data, trigger the execution of the step of sending the first voice data to the cloud server that has established a communication connection with the control device based on the first transmission protocol and subsequent steps;

[0027] When the audio data includes the preset keyword corresponding to the second voice data, trigger the execution of the step of sending the second voice data to the cloud server based on the first transmission protocol and subsequent steps.

[0028] In a possible implementation manner, the first transmission protocol is the websocket protocol, and the second transmission protocol is the HTTPS protocol.

[0029] In a possible implementation manner, the control device is further configured to:

[0030] Send the processed first image data to the LED dot matrix screen for display based on the SPI protocol.

[0031] According to another aspect of the present disclosure, a method for generating an LED dot matrix image is provided, which is used in a control device. The control device is communicatively connected to an LED dot matrix screen and is connected to a voice collection device; the method includes:

[0032] Obtain the first voice data collected by the voice collection device, where the first voice data is used to describe the LED dot matrix image that needs to be displayed on the LED dot matrix screen;

[0033] Send the first voice data to the cloud server that has established a communication connection with the control device based on the first transmission protocol, so that when the cloud server receives the first voice data based on the first transmission protocol, call the first conversion model to convert the first voice data into first text data; and send the first text data to the control device based on the first transmission protocol;

[0034] When receiving the first text data based on the first transmission protocol, send the first text data to the cloud server based on the second transmission protocol, so that when the cloud server receives the first text data based on the second transmission protocol, sequentially call the second conversion model and the third conversion model to convert the first text data into the first text-to-image description text and then generate the first image data; send the first image data to the control device based on the second transmission protocol;

[0035] When the first image data is received based on the second transmission protocol, the first image data is processed to obtain processed first image data suitable for display on the LED dot matrix screen; the processed first image data is sent to the LED dot matrix screen for display.

[0036] According to another aspect of the present disclosure, a method for generating an LED dot matrix image for a cloud server is provided. The method includes:

[0037] When the first voice data sent by the control device is received based on the first transmission protocol, a first conversion model is called to convert the first voice data into first text data; the control device is communicatively connected to the LED dot matrix screen and is provided with a voice collection device; the first voice data is collected by the voice collection device and is used to describe the LED dot matrix image that needs to be displayed on the LED dot matrix screen.

[0038] The first text data is sent to the control device based on the first transmission protocol, so that the control device is further configured to, when the first text data is received based on the first transmission protocol, send the first text data to the cloud server based on the second transmission protocol.

[0039] When the first text data is received based on the second transmission protocol, a second conversion model and a third conversion model are sequentially called to convert the first text data into a first text-to-image description text and then generate first image data.

[0040] The first image data is sent to the control device based on the second transmission protocol, so that the control device, when the first image data is received based on the second transmission protocol, processes the first image data to obtain processed first image data suitable for display on the LED dot matrix screen; the processed first image data is sent to the LED dot matrix screen for display.

[0041] According to another aspect of the present disclosure, an apparatus for generating an LED dot matrix image is provided, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.

[0042] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0043] According to another aspect of the present disclosure, there is provided a computer program product including a computer program or a non-volatile computer-readable storage medium carrying the computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0044] Obtain first voice data collected by a voice collection device through a control device; send the first voice data to a cloud server that has established a communication connection with the control device based on a first transmission protocol; the cloud server is configured to: when receiving the first voice data based on the first transmission protocol, call a first conversion model to convert the first voice data into first text data; and send the first text data to the control device based on the first transmission protocol; when the control device receives the first text data based on the first transmission protocol, send the first text data to the cloud server based on a second transmission protocol; when the cloud server receives the first text data based on the second transmission protocol, sequentially call a second conversion model and a third conversion model to convert the first text data into a first text-to-image description text and then generate first image data; send the first image data to the control device based on the second transmission protocol; when the control device receives the first image data based on the second transmission protocol, process the first image data to obtain processed first image data suitable for display on an LED dot matrix screen; send the processed first image data to the LED dot matrix screen for display; it is possible to generate an LED dot matrix image that conforms to the image described by the first voice data input by the user in real time, which can improve the generation efficiency of the LED dot matrix image and provide a more convenient way to interact with the LED dot matrix screen, and can avoid the problems of the single interaction method of the existing LED dot matrix screen and the low efficiency of generating the LED dot matrix image.

[0045] Meanwhile, the cloud server can assist the control device in generating the first image data described by the first voice data through the first conversion model, the second conversion model, and the third conversion model running internally, which can share the computing tasks of the control device and at the same time improve the generation efficiency of the first image data, thereby further improving the generation efficiency of the LED dot matrix image.

[0046] In addition, the control device interacts with the cloud server based on the first transmission protocol to trigger the cloud server to call the first conversion model, and interacts with the cloud server based on the second transmission protocol to trigger the cloud server to call the second conversion model. Since there are multiple conversion models running in the cloud server, at this time, associating the first transmission protocol with the call of the first conversion model and the second transmission protocol with the call of the second conversion model can avoid the problem of miscalling the conversion model, and at the same time, it can ensure that the transmission protocol is adapted to the data transmission scenario.

[0047] Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The drawings included in and constituting a part of the specification, together with the specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure.

[0049] Figure 1 Schematic structural diagram of a system for generating an LED dot matrix image according to an embodiment of the present disclosure;

[0050] Figure 2 Flowchart of a method for generating an LED dot matrix image according to an embodiment of the present disclosure;

[0051] Figure 3 Flowchart of a method for generating an LED dot matrix image according to another embodiment of the present disclosure;

[0052] Figure 4 Flowchart of a method for generating an LED dot matrix image according to another embodiment of the present disclosure;

[0053] Figure 5 Flowchart of a method for generating an LED dot matrix image according to another embodiment of the present disclosure;

[0054] Figure 6 Flowchart of a method for generating an LED dot matrix image according to another embodiment of the present disclosure;

[0055] Figure 7 Block diagram of an apparatus for generating an LED dot matrix image according to an embodiment of the present disclosure;

[0056] Figure 8 Block diagram of an apparatus for generating an LED dot matrix image according to another embodiment of the present disclosure;

[0057] Figure 9 Block diagram of an apparatus for generating an LED dot matrix image according to another embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the drawings. Identical reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.

[0059] As used herein, the terms "comprising," "including," "having," or variations thereof are open-ended and include one or more stated features, wholes, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, wholes, elements, steps, components, functions, or groups thereof.

[0060] When an element is referred to as being "connected", "coupled", "responsive" or variations thereof to another element, it can be directly connected, coupled or responsive to the other element, or intervening elements may be present.

[0061] Although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Thus, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.

[0062] As used herein, the term "exemplary" means "serving as an example, instance, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments.

[0063] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0064] Figure 1 The structural schematic diagram of a system for generating an LED dot matrix image according to an embodiment of the present disclosure is shown. As Figure 1 shown, the system includes: a control device 110 and a cloud server 120.

[0065] The control device 110 is communicatively connected to the LED dot matrix screen 130, and the control device 110 is connected to a voice collection device 140.

[0066] Optionally, the control device 110 may be implemented in the same device as the LED dot matrix screen 130. For example, the control device 110 is a control chip connected to the LED dot matrix screen 130; or, the control device 110 may also be implemented in a different device from the LED dot matrix screen 130. For example, the control device 110 is implemented in a user terminal independent of the LED dot matrix screen 130, and the user terminal is communicatively connected to the LED dot matrix screen 130 based on a wired or wireless manner. The user terminal includes, but is not limited to, electronic devices with processing capabilities such as computers, tablets, laptops, etc. The implementation manner of the user terminal is not limited in this embodiment. Figure 1Taking the central control device 110 as an ARM processor implemented in the same device as the LED dot matrix screen 130, specifically, the ARM Cortex-M33 processor as an example for illustration, using the ARM Cortex-M33 processor chip for image processing and dot matrix conversion can achieve intelligent display with low power consumption and high efficiency. In actual implementation, the control device 110 can also be implemented in other ways, which are not listed one by one in this embodiment.

[0067] In this embodiment, the voice collection device 140 connected to the control device 110 can be a single microphone, or can also be a microphone array composed of multiple microphones. This embodiment does not limit the implementation manner of the voice collection device 140. Optionally, the voice collection device 140 can be implemented in the same device as the LED dot matrix screen 130, or can also be in a user terminal independent of the LED dot matrix screen 130.

[0068] The control device 110 is used to implement voice interaction between the user and the LED dot matrix screen 130 through the voice collection device 140. In other words, in this embodiment, the control device 110 supports the user to control the LED dot matrix image displayed on the LED dot matrix screen 130 by voice, providing a more convenient interaction method, and at the same time can also improve the generation efficiency of the LED dot matrix image.

[0069] In addition, due to the limited computing power of the control device 110, in this embodiment, the function of generating image data based on voice data is implemented through the cloud server 120, which can share the computing tasks of the control device 110 and at the same time can also improve the generation efficiency of the image data.

[0070] Among them, a first conversion model, a second conversion model, and a third conversion model are running in the cloud server 120. The first conversion model is used to convert voice data into text data; the second conversion model is used to convert the text data output by the first conversion model into text description for text-to-image generation; the third conversion model is used to generate image data matching the text description for text-to-image generation output by the second conversion model.

[0071] Exemplarily, the first conversion model, the second conversion model, and the third conversion model are all open-source models based on neural network models and can be directly applied. Exemplarily, the first conversion model is an Automatic Speech Recognition (ASR) model, such as OpenAI's Whisper model; the second conversion model is the DeepSeek-R1 model; the third conversion model is the Stable Diffusion model. In actual implementation, the first conversion model, the second conversion model, and the third conversion model can also be implemented as other neural network models with corresponding functions. This embodiment does not limit the implementation manners of the first conversion model, the second conversion model, and the third conversion model.

[0072] Exemplarily, referring to Figure 2 , the process in which the cloud server auxiliary control device realizes voice interaction between the user and the LED dot matrix screen through the voice acquisition device includes the following steps:

[0073] Step 201, the control device acquires the first voice data collected by the voice acquisition device, and the first voice data is used to describe the LED dot matrix image to be displayed on the LED dot matrix screen.

[0074] Optionally, the control device controls the voice acquisition device to collect audio data in real time; or, the control device controls the voice acquisition device to collect audio data when it receives a voice input instruction, and controls the voice acquisition device to stop collecting audio data when it receives a stop input instruction. This embodiment does not limit the manner in which the control device controls the voice acquisition device to collect audio data.

[0075] Among them, the acquisition manners of the voice input instruction and the stop input instruction include but are not limited to the following several types:

[0076] The first type: The control device is provided with a voice input control. When a trigger operation acting on the voice input control is received, a voice input instruction is generated; when the trigger operation acting on the voice input control stops executing, a stop input instruction is generated. Among them, the voice input control can be a mechanical button or a virtual control displayed through a touch display screen. This embodiment does not limit the implementation manner of the voice input control.

[0077] The second type: The control device supports communication with a remote control device and receives the voice input instruction and the stop input instruction sent by the remote control device. Among them, the remote control device includes but is not limited to: a remote controller, a user terminal, etc. This embodiment does not limit the implementation manner of the remote control device.

[0078] Optionally, since the audio data collected by the voice acquisition device may be noise, based on this, in order to avoid wasting the computing resources of the cloud server, in this embodiment, the control device can also be used to: determine whether the audio data collected by the voice acquisition device includes a preset keyword to determine whether the audio data is target voice data; the target voice data includes the first voice data and the second voice data described below; in the case where the audio data includes the preset keyword corresponding to the first voice data, it is determined that the first voice data collected by the voice acquisition device is obtained, and the step of triggering and executing sending the first voice data to the cloud server established in communication connection with the control device based on the first transmission protocol and subsequent steps are executed, that is, step 202 is executed.

[0079] At this time, the first voice data includes the preset keyword corresponding to the first voice data and the description information for describing the LED dot matrix image to be displayed on the LED dot matrix screen.

[0080] Exemplarily, determining whether the audio data collected by the voice acquisition device includes a preset keyword includes: converting the audio data into a digital signal and extracting features from the digital signal; comparing the extracted signal features with the template features of the preset keyword; if the signal features include the template features, it is determined that the audio data includes the preset keyword, that is, the audio data is target voice data; if the signal features do not include the template features, it is determined that the audio data does not include the preset keyword, that is, the audio data is not target voice data.

[0081] For example: the preset keyword corresponding to the first voice data is "display". If the audio data collected by the voice acquisition device is "display a smiling face", then the signal features of the audio data include the template features, and it is determined that the audio data is the first voice data.

[0082] In other embodiments, the control device may also not detect whether the audio data includes a preset keyword, but directly use the collected audio data as the first voice data and execute step 202.

[0083] Step 202, the control device sends the first voice data to the cloud server established in communication connection with the control device based on the first transmission protocol.

[0084] The control device and the cloud server are pre-established in communication connection. For example: the control device accesses the Internet based on the WiFi network, so as to access the cloud server based on the Internet. At this time, the control device needs to access a certain WiFi network. Optionally, refer to Figure 1, the control device can configure the WiFi network through Bluetooth. Specifically, after the control device is started, it broadcasts a Bluetooth signal outward through Bluetooth. After the user terminal 150 scans the Bluetooth signal through the Bluetooth scanning function, it sends a Bluetooth connection request to the control device. The control device establishes a communication connection with the user terminal 150 based on the Bluetooth connection request. After the user terminal 150 accesses the WiFi network, it sends the network information of the WiFi network to the control device, and the control device connects to the WiFi network based on the network information. Among them, the network information includes the network name and login password of the WiFi network.

[0085] In other embodiments, in addition to connecting to the WiFi network through Bluetooth, the control device can also connect to the WiFi network through other means. For example, the user terminal sends the network information of the WiFi network to the control device through NFC technology for the control device to connect to the WiFi network based on the network information, etc. Or, the control device can also establish a communication connection with the cloud server in advance based on a cellular network (such as 4G, 5G and other cellular networks). This embodiment does not limit the manner of establishing a communication connection between the control device and the cloud server.

[0086] In this embodiment, the first voice data collected by the control device is streaming data, that is, the first voice data is continuously and real-time generated data, and the first voice data needs to be processed immediately. The websocket protocol is suitable for the streaming data transmission scenario. Based on this, the first transmission protocol is the websocket protocol.

[0087] Step 203, when the cloud server receives the first voice data based on the first transmission protocol, it calls the first conversion model to convert the first voice data into first text data.

[0088] The cloud server listens to the protocol port of the first transmission protocol. If the first voice data is listened to at the protocol port of the first transmission protocol, it calls the first conversion model. After inputting the first voice data into the first conversion model, the first text data corresponding to the first voice data is obtained. Since there are multiple conversion models running in the cloud server, at this time, associating the protocol port of the first transmission protocol with the call of the first conversion model can avoid mis-calling the second conversion model or the third conversion model, and at the same time, it can also be applicable to the streaming data transmission scenario.

[0089] Step 204, the cloud server sends the first text data to the control device based on the first transmission protocol.

[0090] Step 205, when the control device receives the first text data based on the first transmission protocol, it sends the first text data to the cloud server based on the second transmission protocol.

[0091] Since the cloud server needs to generate the first image data expected by the user based on the first text data, this text-to-image task is a one-time task, and the data volume of the first image data is usually large. The HTTPS protocol is suitable for this request-response mode and is suitable for data transmission with a large data volume. Based on this, the second transmission protocol is the HTTPS protocol.

[0092] Optionally, when the control device receives the first text data based on the first transmission protocol, it can also process the first text data into text data suitable for display by the LED dot matrix unit and send the processed text data to the LED dot matrix unit for display. At this time, the user can determine whether there is an error in the conversion content of the cloud server according to the processed text data displayed by the LED dot matrix unit; if there is an error, the user can re-enter the first voice data to trigger the execution of step 201. At this time, the cloud server can re-call the first conversion model to convert the re-entered first voice data; if there is no error, the control device can, after sending the processed text data to the LED dot matrix unit and when the waiting duration reaches the preset waiting duration and no first voice data is received, send the first text data to the cloud server based on the second transmission protocol.

[0093] Among them, the preset waiting duration can be 1s, 2s, etc. This embodiment does not limit the value of the preset waiting duration.

[0094] Step 206, when the cloud server receives the first text data based on the second transmission protocol, it sequentially calls the second conversion model and the third conversion model to convert the first text data into the first text-to-image description text and then generate the first image data.

[0095] The cloud server monitors the protocol port of the second transmission protocol. If the first text data is monitored at the protocol port of the second transmission protocol, it calls the second conversion model. After inputting the first text data into the second conversion model, it obtains the first text-to-image description text (or text-to-image Prompt) corresponding to the first text data. After obtaining the first text-to-image description text, it calls the third conversion model and inputs the first text-to-image description text into the third conversion model to obtain the first image data corresponding to the first text-to-image description text.

[0096] In this embodiment, through the second conversion model, the first text data can be enriched and supplemented while ensuring the user's intention, so as to obtain an image that better meets the user's expectations. For example: the first text data is "display a smiling face", and the second conversion model can combine the user's preferred style (such as cartoon style) to obtain the first text-to-image description text as "a smiling face in cartoon style".

[0097] Step 207, the cloud server sends the first image data to the control device based on the second transmission protocol.

[0098] Step 208, when the control device receives the first image data based on the second transmission protocol, it processes the first image data to obtain the processed first image data suitable for display on the LED dot matrix screen.

[0099] In this embodiment, the control device processes the first image data to obtain the processed first image data, including the following steps:

[0100] Step 1, parse the first image data;

[0101] Exemplarily, parsing the first image data includes: performing data verification on the first image data; after the data verification passes, allocate a buffer in the control device according to the image size indicated by the packet header information of the first image data to store the pixel values of the first image data. Among them, data verification includes but is not limited to: checking the integrity of the first image data.

[0102] Step 2, scale the first image data so that the image resolution of the scaled image data is consistent with the resolution of the LED dot matrix screen;

[0103] In one example, scaling the first image data includes: obtaining the screen size of the LED dot matrix screen and the image size of the first image data; determining the scaling ratio based on the image size and the screen size; using bilinear interpolation based on the scaling ratio to calculate the pixel value mapping of each pixel position in the first image data to the pixel value of the screen position in the LED dot matrix screen to obtain the scaled image data.

[0104] Among them, determining the scaling ratio based on the image size and the screen size can be expressed by the following formula:

[0105]

[0106] Among them, S x represents the scaling ratio of the image width in the image size; W src represents the image width in the image size; W dst represents the screen width in the screen size; S y represents the scaling ratio of the image height in the image size; H src represents the image height in the image size; H dst represents the screen height in the screen size.

[0107] Exemplarily, using bilinear interpolation based on the scaling ratio to calculate the pixel value mapping of each pixel position in the first image data to the pixel value of the screen position in the LED dot matrix screen to obtain the scaled image data includes:

[0108] For each pixel position, determine the floating-point coordinates of that pixel position based on the scaling factor; for example: for the screen position in the i-th row and j-th column, the corresponding floating-point coordinates (or theoretical coordinates) after mapping this screen position back to the first image data are x = i × S x , y = j × S y ;

[0109] Determine four pixel positions Q around the floating-point coordinates 11 (x0, y0), Q 12 (x0, y1), Q 21 (x0 + 1, y0), Q 22 (x0 + 1, y0 + 1), where

[0110] Perform bilinear interpolation on the pixel values of the four pixel positions to obtain the pixel value of the screen position in the i-th row and j-th column; both i and j are positive integers starting from 1 and taking values in sequence; among them, performing bilinear interpolation on the pixel values of the four pixel positions can be expressed by the following formula:

[0111] T dst = (1 - dx)(1 - dy)T 11 + dx(1 - dy)T 21 + (1 - dx)dyT 12 +

[0112] dxdyT 22 ;

[0113] dx = x - x0; dy = y - y0;

[0114] Among them, T dst represents the pixel value of each color channel in the color space of the first image data at the screen position in the i-th row and j-th column. For example: if the color space of the first image data is the RGB space, then T dst represents the color value of the R channel, the color value of the G channel, and the color value of the B channel at the screen position in the i-th row and j-th column respectively; T 11 represents the pixel value of the corresponding color channel at the pixel position Q 11 ; T 21 represents the pixel value of the corresponding color channel at the pixel position Q 21 ; T 12 represents the pixel value of the corresponding color channel at the pixel position Q 12 ; T 22 represents the pixel value of the corresponding color channel at the pixel position Q 22 ;

[0115] In other embodiments, instead of using the bilinear interpolation method, other methods can be used to determine the pixel values at each screen position. For example, the pixel values at the screen positions can be determined through a pre-established mapping relationship between the pixel positions and the screen positions. This embodiment does not limit the method for determining the pixel values at the screen positions.

[0116] Step 3: Convert the pixel values at each screen position into a binary bitmap for driving the LED dot matrix screen to obtain the processed first image data.

[0117] Optionally, if the LED dot matrix screen is a color LED screen, converting the pixel values at each screen position into a binary bitmap for driving the LED dot matrix screen includes: for the pixel values of each color channel at each screen position, extracting each bit of the pixel value represented in binary to form a bit plane, and obtaining a binary matrix formed by the binary values of the bit plane for this bit.

[0118] Alternatively, if the LED dot matrix screen is a monochrome LED screen, converting the pixel values at each screen position into a binary bitmap for driving the LED dot matrix screen includes: converting the color space of the pixel value at this screen position into a monochrome space to obtain the converted pixel value; compressing the pixel values of each row of screen positions or the converted pixel values of each column of screen positions into a byte stream with every 8 pixels as 1 byte to obtain a binary bit stream.

[0119] Exemplarily, converting the color space of the pixel value at this screen position into a monochrome space to obtain the converted pixel value includes: for each screen position, performing a weighted sum of the pixel values of this screen position in each color channel to obtain the grayscale value of this pixel position; if the grayscale value is greater than or equal to a preset monochrome threshold, determining that the converted pixel value at this screen position is 1; if the grayscale value is less than the preset monochrome threshold, determining that the converted pixel value at this screen position is 0.

[0120] For example: the color space of the pixel value at each screen position is in the RGB565 format, that is, the color space of this pixel value is the RGB color space, and the 16-bit pixel value of each pixel value is divided into 3 parts, where the pixel value of red (R) occupies 5 bits, the pixel value of green (G) occupies 6 bits, and the pixel value of blue (B) occupies 5 bits. At this time, convert the pixel value at this screen position from the RGB color space to the monochrome space according to the above process.

[0121] Step 209: The control device sends the processed first image data to the LED dot matrix screen for display.

[0122] Exemplarily, the processed first image data is sent to the LED dot matrix display based on the SPI protocol. Since the SPI protocol supports a relatively high clock frequency, the control device can send the processed first image data to the LED dot matrix display in the form of a high-speed data stream based on the SPI protocol for the LED dot matrix display to display the LED dot matrix image.

[0123] In summary, the system for generating an LED dot matrix image provided in this embodiment obtains the first voice data collected by the voice collection device through the control device; sends the first voice data to the cloud server that has established a communication connection with the control device based on the first transmission protocol; the cloud server is used for: when receiving the first voice data based on the first transmission protocol, calling the first conversion model to convert the first voice data into the first text data; and sending the first text data to the control device based on the first transmission protocol; when the control device receives the first text data based on the first transmission protocol, sending the first text data to the cloud server based on the second transmission protocol; when the cloud server receives the first text data based on the second transmission protocol, sequentially calling the second conversion model and the third conversion model to convert the first text data into the first text-to-image description text and then generate the first image data; sending the first image data to the control device based on the second transmission protocol; when the control device receives the first image data based on the second transmission protocol, processing the first image data to obtain the processed first image data suitable for display on the LED dot matrix display; sending the processed first image data to the LED dot matrix display for display; can generate an LED dot matrix image that conforms to the image described by the first voice data input by the user in real time, can improve the generation efficiency of the LED dot matrix image, and provides a more convenient way to interact with the LED dot matrix display, and can avoid the problems of single interaction method of the existing LED dot matrix display and low efficiency of generating the LED dot matrix image.

[0124] At the same time, the cloud server can assist the control device in generating the first image data described by the first voice data through the first conversion model, the second conversion model, and the third conversion model running internally, which can share the computing tasks of the control device and at the same time improve the generation efficiency of the first image data, thereby further improving the generation efficiency of the LED dot matrix image.

[0125] In addition, the control device interacts with the cloud server based on the first transmission protocol to trigger the cloud server to call the first conversion model, and interacts with the cloud server based on the second transmission protocol to trigger the cloud server to call the second conversion model. Since there are multiple conversion models running in the cloud server, at this time, associating the first transmission protocol with the call of the first conversion model and the second transmission protocol with the call of the second conversion model can avoid the problem of miscalling the conversion model, and at the same time, it can ensure that the transmission protocol is adapted to the data transmission scenario.

[0126] Optionally, in the above embodiments, steps 201, 202, 205, 208, and 209 can be separately implemented as embodiments on the control device side, and steps 203, 204, 206, and 207 can be separately implemented as embodiments on the control device side.

[0127] Optionally, since the user may also need to modify the LED dot matrix image displayed on the LED dot matrix screen, based on this, after step 209, referring to Figure 3 , the process in which the cloud server assists the control device to implement voice interaction between the user and the LED dot matrix screen through the voice collection device includes the following steps:

[0128] Step 301, after the control device sends the processed image data to the LED dot matrix screen for display, if the second voice data is received, the second voice data is sent to the cloud server based on the first transmission protocol.

[0129] Among them, the second voice data is used to describe the content that needs to be modified in the LED dot matrix image currently displayed on the LED dot matrix screen.

[0130] The second voice data is data collected by the control device in real time through the voice collection device.

[0131] Optionally, the manner in which the control device receives the second voice data includes: after sending the processed image data to the LED dot matrix screen for display, outputting an image modification prompt, which is used to prompt the user to determine whether to modify the LED dot matrix image currently displayed on the LED dot matrix screen, and outputting the second voice data when the LED dot matrix image needs to be modified; correspondingly, the control device obtains the second voice data.

[0132] Among them, the image modification prompt includes but is not limited to: displaying a text prompt through the LED dot matrix screen, or outputting an audio prompt, etc. The present embodiment does not limit the output form of the image modification prompt.

[0133] In a possible implementation manner, if the duration of the image modification prompt output by the control device reaches the preset waiting duration and the second voice data is not received, it is determined that the user does not need to modify the LED dot matrix image currently displayed on the LED dot matrix screen, and the output of the image modification prompt is stopped. Or, if the duration of the image modification prompt output by the control device does not reach the preset waiting duration and the second voice data is received, it is determined that the user needs to modify the LED dot matrix image currently displayed on the LED dot matrix screen, and the output of the image modification prompt is stopped.

[0134] Optionally, since the audio data received by the control device may also be noise after the processed image data is sent to the LED dot matrix screen for display, based on this, in order to avoid wasting the computing resources of the cloud server, in this embodiment, before the control device sends the second voice data to the cloud server based on the first transmission protocol, it is further configured to: determine whether the audio data collected by the voice collection device includes a preset keyword to determine whether the audio data is target voice data; at this time, the target voice data includes the second voice data; in the case that the voice data includes the preset keyword corresponding to the second voice data, trigger the execution of the step of sending the second voice data to the cloud server based on the first transmission protocol and subsequent steps.

[0135] For the relevant description of determining whether the audio data collected by the voice collection device includes a preset keyword, see the above embodiment, and this embodiment will not be elaborated here.

[0136] For example: the preset keyword corresponding to the second voice data is "modify". If the audio data collected by the voice collection device is "make the eyes bigger", then the signal feature of the audio data includes the template feature, and it is determined that the audio data is the second voice data.

[0137] Step 302, when the cloud server receives the second voice data based on the first transmission protocol, call the first conversion model to convert the second voice data into second text data; and send the second text data to the control device based on the first transmission protocol.

[0138] Step 303, when the control device receives the second text data based on the first transmission protocol, send the second text data to the cloud server based on the second transmission protocol.

[0139] Step 304, when the cloud server receives the second text data based on the second transmission protocol, call the second conversion model and the third conversion model in sequence to convert the second text data into a second text-to-image description text and then generate second image data; send the second image data to the control device based on the second transmission protocol.

[0140] For the relevant description of Steps 302-304, see Steps 202-207. Just update the first voice data, the first text data, the first text-to-image description text, and the first image data in Steps 202-207 to the second voice data, the second text data, the second text-to-image description text, and the second image data respectively.

[0141] Step 305, when the control device receives the second image data based on the second transmission protocol, process the second image data; send the processed second image data to the LED dot matrix screen to modify the LED dot matrix image currently displayed on the LED dot matrix screen.

[0142] In one example, the processing in step 305 is the same as that in step 208. At this time, the control device maps the second image data as a whole into image data suitable for display on the LED dot matrix screen for the LED dot matrix screen to display.

[0143] In another example, when the second voice data indicates to modify a local area in the LED dot matrix image currently displayed on the LED dot matrix screen, the control device may only map the partial image data corresponding to the local area into image data suitable for display on the LED dot matrix screen. In this way, it can be ensured that other parts of the modified LED dot matrix image do not need to be changed, ensuring the display effect of the LED dot matrix image. At the same time, the data transmission volume is reduced and communication resources are saved.

[0144] At this time, the cloud server is further configured to: after successively invoking the second conversion model and the third conversion model to generate the second image data, obtain the local pixel positions of the local area in the second image data, and send the local pixel positions to the control device.

[0145] Among them, the local pixel positions are used to indicate the positions of the local areas in the second image data. Optionally, the second image data sent by the cloud server to the control device according to the second transmission protocol may only include the image data of the local pixel positions; or it may also be all of the second image data.

[0146] The methods for the cloud server to obtain the local pixel positions include but are not limited to:

[0147] The first method: Compare the second image data with the first image data pixel by pixel to mark the different areas; use connected component analysis to merge the discrete different points into continuous areas to obtain the local pixel positions of the continuous areas.

[0148] The second method: Invoke a pre-created target detection model to detect the modification target indicated by the second voice data in the second image data to obtain the local pixel positions of the target in the second image data.

[0149] The control device is further configured to: map the image content of the local pixel positions in the second image data to the corresponding screen positions of the local areas on the LED dot matrix screen to obtain the processed second image data; send the processed second image data to the LED dot matrix screen to modify the LED dot matrix image currently displayed on the LED dot matrix screen.

[0150] Map the image content at the local pixel positions to the corresponding screen positions in the LED dot matrix screen, including: obtaining a scaling ratio; mapping the local pixel positions to local screen positions in the LED dot matrix screen according to the scaling ratio; using bilinear interpolation to calculate the pixel values at each local screen position to obtain the pixel content at each local screen position.

[0151] Since the second image data and the first image data are generated based on the same third conversion model, the image size of the second image data is equal to the image size of the first image data. Accordingly, the scaling ratio is the scaling ratio determined based on the image size of the first image data and the screen size. For specific reference, see the above system embodiment.

[0152] Map the local pixel position (x p , y p ) to the local screen position (x led , y led ) in the LED dot matrix screen according to the scaling ratio, which can be represented by the following formula:

[0153] x led = x p / S x ;

[0154] y led = y p / S y ;

[0155] where S x represents the scaling ratio of the image width in the image size, and S y represents the scaling ratio of the image height in the image size.

[0156] For the relevant description of using bilinear interpolation to calculate the pixel values at each local screen position, see the content of step 2 in the above system embodiment, which will not be elaborated here in this embodiment.

[0157] Optionally, if the processed second image data only includes the binary bitmaps at the local screen positions, the control device also needs to send the local screen positions to the LED dot matrix screen to instruct the LED dot matrix screen to control the states of the LEDs at the local screen positions according to the binary bitmaps; or, if the processed second image data includes the binary bitmaps at all screen positions, the control device sends the binary bitmaps at all screen positions to the LED dot matrix screen to instruct the LED dot matrix screen to control the states of the LEDs at each screen position according to the binary bitmaps.

[0158] Optionally, after the LED dot matrix screen displays the modified LED dot matrix image, step 301 can be executed again to modify the LED dot matrix image again.

[0159] In this embodiment, after the processed image data is sent to the LED dot matrix screen for display, the second voice data is received to modify the LED dot matrix image displayed on the LED dot matrix screen, which can improve the efficiency of modifying the LED dot matrix image.

[0160] In addition, by only mapping the local pixel positions in the local area, the control device can reduce the computational task volume of the control device and further improve the modification efficiency of the LED dot matrix image.

[0161] To more clearly understand the process of generating the LED dot matrix image provided in this application, an example of this process will be described below. Refer to Figure 4 , this process includes the following steps:

[0162] Step 41, establish a communication connection between the control device and the cloud server;

[0163] Step 42, the control device receives the voice data input by the user and sends the voice data to the cloud server;

[0164] Among them, the voice data can be the first voice data or the second voice data;

[0165] Step 43, the cloud server receives the voice data, calls the first conversion model to convert the voice data into text data, and sends the text data to the control device;

[0166] Step 44, the control device receives the text data and sends the text data to the cloud server;

[0167] Step 45, the cloud server receives the text data, calls the second conversion model to convert the text data into a text-to-image description text, and calls the third conversion model to generate image data based on the text-to-image description text, and returns the image data to the control device;

[0168] Step 46, the control device receives the image data and converts the image data into processed image data suitable for display on the LED dot matrix screen;

[0169] Step 47, the control device sends the processed image data to the LED dot matrix screen for display;

[0170] Step 48, the control device determines whether to modify the image data; if so, execute Step 42; if not, the process ends.

[0171] The ways for the control device to determine whether to modify the image data include: determining whether voice data (i.e., the second voice data) is received within a preset waiting duration; if so, determining to modify the image data; if not, determining not to modify the image data. Alternatively, determining whether a modification instruction sent by a user terminal communicatively connected to the control device is received; if so, determining to modify the image data; if not, determining not to modify the image data. In other embodiments, the ways to determine whether to modify the image data can also be other ways, which are not limited in this embodiment.

[0172] In this embodiment, the automatic conversion from voice to image is realized through a multimodal AI model, endowing the LED dot matrix screen with intelligent capabilities; users can modify the content of the LED dot matrix image through voice description, improving the operation convenience.

[0173] Figure 5 The flowchart showing a method for generating an LED dot matrix image according to an embodiment of the present disclosure is as follows. As Figure 5 shown, this embodiment takes the method being used in Figure 1 the control device shown as an example for illustration. The method includes:

[0174] Step 501, obtain first voice data collected by a voice acquisition device, where the first voice data is used to describe the LED dot matrix image that needs to be displayed on the LED dot matrix screen;

[0175] Step 502, send the first voice data to a cloud server communicatively connected to the control device based on a first transmission protocol, so that when the cloud server receives the first voice data based on the first transmission protocol, it calls a first conversion model to convert the first voice data into first text data; and send the first text data to the control device based on the first transmission protocol;

[0176] Step 503, when the first text data is received based on the first transmission protocol, send the first text data to the cloud server based on a second transmission protocol, so that when the cloud server receives the first text data based on the second transmission protocol, it sequentially calls a second conversion model and a third conversion model to convert the first text data into a first text-to-image description text and then generate first image data; send the first image data to the control device based on the second transmission protocol;

[0177] Step 504, when the first image data is received based on the second transmission protocol, process the first image data to obtain processed first image data suitable for display on the LED dot matrix screen; send the processed first image data to the LED dot matrix screen for display.

[0178] For relevant details, refer to the above system embodiment.

[0179] In this embodiment, an LED dot matrix image that conforms to the image described by the first voice data input by the user can be generated in real time, which can improve the generation efficiency of the LED dot matrix image and provide a more convenient way to interact with the LED dot matrix screen, and can avoid the problems of the single interaction method of the existing LED dot matrix screen and the low efficiency of generating the LED dot matrix image.

[0180] Figure 6 The flowchart showing the method for generating an LED dot matrix image according to an embodiment of the present disclosure is as follows. Figure 6 As shown, this embodiment takes the method being used in Figure 1 the cloud server shown as an example for illustration. The method includes:

[0181] Step 601, when receiving the first voice data sent by the control device based on the first transmission protocol, call the first conversion model to convert the first voice data into first text data; the control device is communicatively connected to the LED dot matrix screen and is provided with a voice collection device; the first voice data is collected by the voice collection device and is used to describe the LED dot matrix image that needs to be displayed on the LED dot matrix screen;

[0182] Step 602, send the first text data to the control device based on the first transmission protocol, so that the control device is further used to, when receiving the first text data based on the first transmission protocol, send the first text data to the cloud server based on the second transmission protocol;

[0183] Step 603, when receiving the first text data based on the second transmission protocol, call the second conversion model and the third conversion model in sequence to convert the first text data into a first text-to-image description text and then generate first image data; send the first image data to the control device based on the second transmission protocol, so that the control device, when receiving the first image data based on the second transmission protocol, processes the first image data to obtain the processed first image data suitable for display on the LED dot matrix screen; send the processed first image data to the LED dot matrix screen for display.

[0184] For relevant details, refer to the above system embodiment.

[0185] In this embodiment, an LED dot matrix image that conforms to the image described by the first voice data input by the user can be generated in real time, which can improve the generation efficiency of the LED dot matrix image and provide a more convenient way to interact with the LED dot matrix screen, and can avoid the problems of the single interaction method of the existing LED dot matrix screen and the low efficiency of generating the LED dot matrix image.

[0186] Figure 7 The block diagram showing the device for generating an LED dot matrix image according to an embodiment of the present disclosure is as follows.Figure 7 As shown in Figure 7 , in a control device, the device includes: a voice acquisition module 710, a first transmission module 720, a second transmission module 730, and an image processing module 740.

[0187] The voice acquisition module 710 is configured to acquire first voice data collected by the voice collection device, and the first voice data is used to describe an LED dot matrix image that needs to be displayed on the LED dot matrix screen.

[0188] The first transmission module 720 is configured to send the first voice data to a cloud server that has established a communication connection with the control device based on a first transmission protocol. For the cloud server, when receiving the first voice data based on the first transmission protocol, the cloud server is configured to call a first conversion model to convert the first voice data into first text data; and send the first text data to the control device based on the first transmission protocol.

[0189] The second transmission module 730 is configured to, when receiving the first text data based on the first transmission protocol, send the first text data to the cloud server based on a second transmission protocol. For the cloud server, when receiving the first text data based on the second transmission protocol, the cloud server is configured to sequentially call a second conversion model and a third conversion model to convert the first text data into a first text-to-image description text and then generate first image data; and send the first image data to the control device based on the second transmission protocol.

[0190] The image processing module 740 is configured to, when receiving the first image data based on the second transmission protocol, process the first image data to obtain processed first image data suitable for display on the LED dot matrix screen; and send the processed first image data to the LED dot matrix screen for display.

[0191] For details of the related description, refer to the above embodiments.

[0192] Figure 8 The block diagram of a device for generating an LED dot matrix image according to an embodiment of the present disclosure is shown. As Figure 8 shown, in a cloud server, the device includes: a first conversion module 810, a first transmission module 820, a second conversion module 830, and a second transmission module 840.

[0193] The first conversion module 810 is configured to, when receiving the first voice data sent by the control device based on the first transmission protocol, call the first conversion model to convert the first voice data into first text data; the control device is communicatively connected to the LED dot matrix screen and is provided with a voice collection device; the first voice data is collected by the voice collection device and is used to describe the LED dot matrix image to be displayed on the LED dot matrix screen.

[0194] The first transmission module 820 is configured to send the first text data to the control device based on the first transmission protocol, so that the control device is further configured to, when receiving the first text data based on the first transmission protocol, send the first text data to the cloud server based on the second transmission protocol.

[0195] The second conversion module 830 is configured to, when receiving the first text data based on the second transmission protocol, sequentially call the second conversion model and the third conversion model to convert the first text data into the first text-to-image description text and then generate the first image data.

[0196] The second transmission module 840 is configured to send the first image data to the control device based on the second transmission protocol, so that the control device, when receiving the first image data based on the second transmission protocol, processes the first image data to obtain the processed first image data suitable for display on the LED dot matrix screen; and sends the processed first image data to the LED dot matrix screen for display.

[0197] For details of the relevant descriptions, please refer to the above embodiments.

[0198] In some embodiments, the functions or modules included in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the method embodiments above. The specific implementation can refer to the descriptions of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0199] The embodiments of the present disclosure further provide a device for generating an LED dot matrix image, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.

[0200] The embodiments of the present disclosure further provide a non-volatile computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0201] The embodiments of the present disclosure further provide a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0202] Figure 9 is a block diagram of an apparatus 1900 for generating an LED dot matrix image shown according to an exemplary embodiment. For example, the apparatus 1900 may be provided as a server or a terminal device. Referring to Figure 9 , the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0203] The apparatus 1900 may further include a power component 1926 configured to perform power management of the apparatus 1900, a wired or wireless network interface 1950 configured to connect the apparatus 1900 to a network, and an input / output interface 1958 (I / O interface). The apparatus 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or the like.

[0204] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the apparatus 1900 to complete the above method.

[0205] A computer-readable storage medium can be a tangible device that can hold and store programs / instructions used by an instruction execution device. A computer-readable storage medium can be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0206] The computer programs (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0207] A computer program (or computer program instructions) for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on a user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0208] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0209] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, so that the computer-readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0210] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0211] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0212] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A system for generating LED dot matrix images, characterized in that: The system includes a control device and a cloud server; The control device is connected to the LED dot matrix screen for communication, and the control device is connected to a voice collection device; the control device is used to: obtain first voice data collected by the voice collection device, the first voice data is used to describe the LED dot matrix image that needs to be displayed by the LED dot matrix screen; send the first voice data to a cloud server that establishes a communication connection with the control device based on a first transmission protocol; The cloud server is used for: upon receiving the first voice data based on the first transmission protocol, calling the first conversion model to convert the first voice data into first text data; and sending the first text data to the control device based on the first transmission protocol; The control device is further configured to: upon receiving the first text data based on the first transmission protocol, send the first text data to the cloud server based on a second transmission protocol; The cloud server is further used for: upon receiving the first text data based on the second transmission protocol, calling the second conversion model and the third conversion model in sequence to convert the first text data into a first text-image description text to generate first image data; and sending the first image data to the control device based on the second transmission protocol; The control device is further used for: in the case of receiving the first image data based on the second transmission protocol, processing the first image data to obtain processed first image data suitable for display on the LED dot matrix screen; The processed first image data is sent to the LED dot matrix screen for display.

2. The system according to claim 1, characterized in that The control device is also used for: after sending the processed image data to the LED dot matrix screen for display, if second voice data is received, sending the second voice data to the cloud server based on the first transmission protocol; wherein the second voice data is used to describe the content that needs to be modified in the LED dot matrix image currently displayed by the LED dot matrix screen; The cloud server is further configured to: upon receiving the second voice data based on the first transmission protocol, call the first conversion model to convert the second voice data into second text data; and send the second text data to the control device based on the first transmission protocol; The control device is further configured to: upon receiving the second text data based on the first transmission protocol, send the second text data to the cloud server based on the second transmission protocol; The cloud server is further used for: upon receiving the second text data based on the second transmission protocol, calling the second conversion model and the third conversion model in sequence to convert the second text data into a second text-image description text to generate second image data; and sending the second image data to the control device based on the second transmission protocol; The control device is also used to: when the second image data is received based on the second transmission protocol, process the second image data; and send the processed second image data to the LED dot matrix screen to modify the LED dot matrix image currently displayed by the LED dot matrix screen.

3. The system according to claim 2, characterized in that The cloud server is further used for: when the second voice data indicates to modify the local area in the LED dot matrix image currently displayed by the LED dot matrix screen, after sequentially calling the second conversion model and the third conversion model to generate the second image data, obtaining the local pixel position of the local area in the second image data, and sending the local pixel position to the control device; The control device is also used to: map the image content of the local pixel position in the second image data to the screen position corresponding to the local area in the LED dot matrix screen to obtain the processed second image data; The processed second image data is sent to the LED dot matrix screen to modify the LED dot matrix image currently displayed by the LED dot matrix screen.

4. The system according to claim 2, characterized in that The control device is also used for: After the processed image data is sent to the LED dot matrix screen for display, an image modification prompt is output, wherein the image modification prompt is used to prompt the user to determine whether the LED dot matrix image currently displayed on the LED dot matrix screen needs to be modified, and the second voice data is output when the LED dot matrix image needs to be modified.

5. The system according to claim 2, characterized in that The control device is also used for: Determining whether the audio data collected by the voice collection device includes a preset keyword to determine whether the audio data is target voice data, wherein the target voice data includes the first voice data and the second voice data; In a case where the audio data includes a preset keyword corresponding to the first voice data, triggering execution of the step of sending the first voice data to a cloud server that has established a communication connection with the control device based on a first transmission protocol and subsequent steps; In a case where the audio data includes a preset keyword corresponding to the second voice data, the step of sending the second voice data to the cloud server based on the first transmission protocol and subsequent steps are triggered.

6. The system according to any one of claims 1 to 5, characterized in that: The first transmission protocol is the websocket protocol, and the second transmission protocol is the HTTPS protocol.

7. The system according to any one of claims 1 to 5, characterized in that: The control device is also used for: The processed first image data is sent to the LED dot matrix screen for display based on the SPI protocol.

8. A method for generating an LED dot matrix image, characterized in that: Used in a control device, the control device is communicatively connected to an LED dot matrix screen and is connected to a voice collection device; The method comprises: Acquire first voice data collected by the voice collection device, where the first voice data is used to describe an LED dot matrix image that needs to be displayed by the LED dot matrix screen; The first voice data is sent to a cloud server that establishes a communication connection with the control device based on a first transmission protocol, so that the cloud server, when receiving the first voice data based on the first transmission protocol, calls a first conversion model to convert the first voice data into first text data; and sends the first text data to the control device based on the first transmission protocol; When the first text data is received based on the first transmission protocol, the first text data is sent to the cloud server based on the second transmission protocol, so that the cloud server, when the first text data is received based on the second transmission protocol, sequentially calls the second conversion model and the third conversion model to convert the first text data into a first text image description text to generate the first image data; and sends the first image data to the control device based on the second transmission protocol; When the first image data is received based on the second transmission protocol, the first image data is processed to obtain processed first image data suitable for display on the LED dot matrix screen; and the processed first image data is sent to the LED dot matrix screen for display.

9. A method for generating an LED dot matrix image, characterized in that: For a cloud server, the method includes: In the case of receiving first voice data sent by the control device based on the first transmission protocol, calling the first conversion model to convert the first voice data into first text data; the control device is connected to the LED dot matrix screen for communication and is provided with a voice collection device; the first voice data is collected by the voice collection device and is used to describe the LED dot matrix image that needs to be displayed by the LED dot matrix screen; sending the first text data to the control device based on the first transmission protocol, so that the control device is further used to send the first text data to the cloud server based on a second transmission protocol when the first text data is received based on the first transmission protocol; When the first text data is received based on the second transmission protocol, the second conversion model and the third conversion model are called in sequence to convert the first text data into a first text-image description text to generate first image data; The first image data is sent to the control device based on the second transmission protocol, so that the control device processes the first image data when receiving the first image data based on the second transmission protocol to obtain the processed first image data suitable for display on the LED dot matrix screen; and the processed first image data is sent to the LED dot matrix screen for display.

10. A device for generating an LED dot matrix image, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to claim 8 or 9.