A data processing method and system

By obtaining and processing the user's voice signals and image files in real time, dynamically adjusting the ambient light intensity, solving the problem that video content cannot automatically adjust the ambient light in the prior art, and improving the user's viewing experience.

CN114650640BActive Publication Date: 2025-06-17BEIJING LEXIN ZHONGLIAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210363673.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-07
Publication Date
2025-06-17
Estimated Expiration
2042-04-07

AI Technical Summary

Technical Problem

The prior art is difficult to automatically adjust the ambient light intensity of video content, resulting in poor user viewing experience.

Method used

By obtaining the user's voice signal in real time, vocal recognition is performed and converted into audio signals, the content of the audio signal is recognized to determine the reference brightness and light adjustment model, the adjustment image group is determined based on the repetitive recognition results of the image file, and input it into the light adjustment model to generate adjustment instructions.

Benefits of technology

It realizes dynamic adjustment of the ambient light intensity based on the video content, improves the user's viewing experience, and builds a dynamic viewing platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114650640B_ABST
    Figure CN114650640B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of smart home, and specifically discloses a data processing method and system. The method includes determining a reference brightness and a corresponding lighting adjustment model, generating a reference adjustment instruction based on the reference brightness, and sending the reference adjustment instruction to an execution end; obtaining an image file based on a play request input by a user, performing repetitive recognition on the image file, and determining an adjustment image group according to the repetitive recognition result; inputting the adjustment image group into the lighting adjustment model to generate a lighting adjustment instruction, and sending the lighting adjustment instruction to the execution end. The present invention first determines the reference brightness, then performs content recognition on the image file to be played based on a preset lighting adjustment model to generate a lighting adjustment instruction, and further adjusts the ambient light intensity based on the lighting adjustment instruction, successfully building a dynamic viewing platform and improving the user's viewing experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart home, and specifically to a data processing method and system. Background Art

[0002] With the development of society, videos have gradually become a mainstream way of entertainment or information transmission. Specifically in practice, it is the movies, TV dramas or short videos that we usually watch. Among them, due to the high production cost of movies, users generally have relatively high requirements for the viewing environment and always hope to obtain better external conditions for watching movies. Among them, the environmental light intensity is a very important influencing factor.

[0003] In the prior art, there are many lights with adjustable brightness, but there is no light that can automatically change with the video content. How to make the external environment better adapt to the video content, build a good private cinema, and improve the user's viewing experience is the technical problem that the technical solution of the present invention wants to solve. Summary of the Invention

[0004] The purpose of the present invention is to provide a data processing method and system to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solutions:

[0006] A data processing method, the method includes:

[0007] Real-time obtain the voice signal sent by the user, and perform voice recognition on the voice signal; when the voice signal is a voice signal, convert the voice signal into an audio signal;

[0008] Perform content recognition on the audio signal, determine the reference brightness and the corresponding light adjustment model, generate a reference adjustment instruction based on the reference brightness, and send the reference adjustment instruction to the execution end;

[0009] Obtain an image file based on the play request input by the user, perform repeatability recognition on the image file, and determine an adjustment image group according to the repeatability recognition result;

[0010] Input the adjustment image group into the light adjustment model, generate a light adjustment instruction, and send the light adjustment instruction to the execution end.

[0011] As a further solution of the present invention: the step of real-time obtaining the voice signal sent by the user and performing voice recognition on the voice signal includes:

[0012] Real-time obtain the voice signal sent by the user, input the voice signal into a trained voice recognition model, and obtain a recognition text;

[0013] Judge whether the voice signal is an intentional signal according to the recognized text;

[0014] When the voice signal is an intentional signal, determine the voice signal information points according to a preset amplitude threshold;

[0015] Judge whether the intentional signal is a human voice signal according to the voice signal information points;

[0016] As a further solution of the present invention: the step of judging whether the intentional signal is a human voice signal according to the voice signal information points includes:

[0017] Obtain the amplitudes of several voice signal information points;

[0018] Obtain the phase differences of several adjacent voice signal information points;

[0019] Calculate the means of the obtained amplitudes and the phase differences respectively, and calculate the variance according to the means;

[0020] Compare the calculated variance with a preset variance threshold, and judge whether the intentional signal is a human voice signal according to the comparison result;

[0021] As a further solution of the present invention: the step of obtaining an image file based on a user input playback request, performing repetitive recognition on the image file, and determining an adjusted image group according to the repetitive recognition result includes:

[0022] Receive a playback request containing a file index input by the user, and extract the image file from a preset file library based on the file index; when the image file does not exist in the preset file library, open a file acquisition port and acquire the image file based on the file acquisition port;

[0023] Extract the image information in the image file to generate an image sequence; the number of the image sequence has a mapping relationship with the time information in the image information;

[0024] Convert the image sequence into a single-value image library according to a preset conversion formula;

[0025] Perform repetitive recognition on the single-value image library, mark the repeated images in the single-value image library based on the repetitive recognition result, and remove the repeated images in the image file based on the mapping relationship to obtain an adjusted image group;

[0026] As a further solution of the present invention: the step of performing repetitive recognition on the single-value image library and marking the repeated images in the single-value image library based on the repetitive recognition result includes:

[0027] Read the single - value images in the single - value image library in sequence, traverse each pixel point in the single - value image, and obtain the value of each pixel point;

[0028] Calculate the eigenvalue of the single - value image based on the values of each pixel point, and obtain an eigenvalue group corresponding to the single - value image library;

[0029] Calculate the offset rate of adjacent eigenvalues in the eigenvalue group in sequence, and compare the offset rate with a preset offset threshold;

[0030] When the offset rate is less than the preset offset threshold, mark the single - value image corresponding to the subsequent eigenvalue as a duplicate image.

[0031] As a further solution of the present invention: the step of inputting the adjustment image group into the lighting adjustment model to generate a lighting adjustment instruction includes:

[0032] Read the eigenvalue group corresponding to the adjustment image group;

[0033] Generate a fitting curve based on the eigenvalue group, input the fitting curve into a trained instruction generation model, and obtain a pre - adjustment instruction;

[0034] Obtain the adjustment image corresponding to the eigenvalue, obtain the time information of the adjustment image, and insert the time information into the pre - adjustment instruction to obtain a lighting adjustment instruction.

[0035] As a further solution of the present invention: the method further includes:

[0036] Receive a brightness setting request sent by the user and open a brightness setting port;

[0037] Receive a touch - screen signal input by the user based on the brightness setting port, and determine a brightness setting value based on the touch - screen signal;

[0038] Replace the reference brightness according to the brightness setting value.

[0039] The technical solution of the present invention also provides a data processing system, and the system includes:

[0040] A voice recognition module, which is used to obtain a voice signal sent by the user in real - time and perform voice recognition on the voice signal; when the voice signal is a voice signal, convert the voice signal into an audio signal;

[0041] A reference adjustment module, which is used to perform content recognition on the audio signal, determine a reference brightness and a corresponding lighting adjustment model, generate a reference adjustment instruction based on the reference brightness, and send the reference adjustment instruction to the execution end;

[0042] The repeated image recognition module obtains video files based on the playback requests input by users, performs repeated recognition on the video files, and determines an adjusted image group according to the results of the repeated recognition;

[0043] The instruction generation module is used to input the adjusted image group into the lighting adjustment model, generate a lighting adjustment instruction, and send the lighting adjustment instruction to the execution end.

[0044] As a further solution of the present invention: the voice recognition module includes:

[0045] The text generation unit is used to obtain the voice signals sent by users in real time, input the voice signals into the trained voice recognition model, and obtain the recognized text;

[0046] The meaning determination unit is used to judge whether the voice signal is an intentional signal according to the recognized text;

[0047] The information point determination unit is used to determine the voice signal information points according to the preset amplitude threshold when the voice signal is an intentional signal;

[0048] The processing execution unit is used to judge whether the intentional signal is a human voice signal according to the voice signal information points.

[0049] As a further solution of the present invention: the processing execution unit includes:

[0050] The amplitude acquisition sub-unit is used to acquire the amplitudes of several voice signal information points;

[0051] The phase difference acquisition sub-unit is used to acquire the phase differences of several adjacent voice signal information points;

[0052] The variance calculation sub-unit is used to calculate the means of the acquired amplitudes and phase differences respectively, and calculate the variance according to the means;

[0053] The comparison sub-unit is used to compare the calculated variance with the preset variance threshold, and judge whether the intentional signal is a human voice signal according to the comparison result.

[0054] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention first determines the reference brightness, then performs content recognition on the video files to be played based on the preset lighting adjustment model, generates a lighting adjustment instruction, and further adjusts the ambient light intensity based on the lighting adjustment instruction, successfully building a dynamic movie-watching platform and improving the movie-watching experience of users. Description of the Drawings

[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for use in the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention.

[0056] Figure 1 It is a flowchart of a data processing method.

[0057] Figure 2 It is the first sub-flowchart of the data processing method.

[0058] Figure 3 It is the second sub-flowchart of the data processing method.

[0059] Figure 4 It is the third sub-flowchart of the data processing method.

[0060] Figure 5 It is the fourth sub-flowchart of the data processing method.

[0061] Figure 6 It is the fifth sub-flowchart of the data processing method.

[0062] Figure 7 It is a block diagram of the composition structure of a data processing system.

[0063] Figure 8 It is a block diagram of the composition structure of the speech recognition module in the data processing system.

[0064] Figure 9 It is a block diagram of the composition structure of the processing execution unit in the speech recognition module. Detailed implementation manners

[0065] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the following further details the present invention in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0066] Embodiment 1

[0067] Figure 1 It is a flowchart of a data processing method. In the embodiments of the present invention, a data processing method, the method includes steps S100 to step S400:

[0068] Step S100: Real-time obtain the voice signal sent by the user, and perform human voice recognition on the voice signal; when the voice signal is a human voice signal, convert the voice signal into an audio signal;

[0069] One of the application scenarios of the technical solution of the present invention is a private cinema, which has a voice control function. When a user enters a certain area, the entire system can be made to enter the preparation stage through voice.

[0070] Step S200: Perform content recognition on the audio signal, determine the reference brightness and the corresponding lighting adjustment model, generate a reference adjustment instruction based on the reference brightness, and send the reference adjustment instruction to the execution end;

[0071] The lighting adjustment model may not be unique. For example, when watching some suspense movies, when encountering darker video parts, some users are more afraid and they hope the environment is brighter to increase a sense of security, while some users are bolder and they hope the environment is darker to increase the sense of immersion. It can be imagined that the adjustment processes of these two lighting adjustment models are opposite. The purpose of step S200 is only to determine the lighting adjustment model and the reference brightness according to the user's voice information. Among them, the reference brightness is the brightness of the lights in the area, and the adjustment process is based on the enhancement or weakening of this reference brightness.

[0072] Step S300: Obtain an image file based on the playback request input by the user, perform repetitive recognition on the image file, and determine an adjustment image group according to the repetitive recognition result;

[0073] Step S300 is the processing process of the image file. When adjusting the brightness based on the image file, if the adjustment is performed in real time, the computing resources consumed will be very high. In fact, the image file is composed of small segments. In each small segment, the change range of the brightness (color value of the image) of the image is not large. Only when switching scenes, the change range of the image will be very large, and correspondingly, the indoor brightness also changes accordingly.

[0074] Step S400: Input the adjustment image group into the lighting adjustment model, generate a lighting adjustment instruction, and send the lighting adjustment instruction to the execution end.

[0075] Step S400 is the specific process of adjusting the ambient brightness. The lighting adjustment instruction refers to enhancement or weakening. The range of enhancement or weakening is approximately around 10% or 20% of the reference brightness. Generally, it will not exceed 50%. If it exceeds 50%, then adjusting the reference brightness is a better solution.

[0076] The architecture of the technical solution of the present invention includes two parties, namely the user terminal and the execution end. Data is transmitted between the user terminal and the execution end through a network, and the connection type of the network is mainly a wireless communication link.

[0077] During the process of using the user terminal, the user can interact with the execution end through the network. The user terminal can be hardware or software. When the user terminal is hardware, the user terminal is a portable device, and the most common portable device is a mobile phone. When the user terminal is software, it can be installed in the above device, and it can be implemented as multiple software or software modules, or it can be implemented as a single software or software module, which is not specifically limited here.

[0078] The execution end can be hardware or software. When the device terminal is hardware, the execution end is a programmable lamp that can receive the display brightness or the corrected brightness and generate a corresponding light source intensity instruction based on the display brightness or the corrected brightness, thereby adjusting the light source intensity. When the execution end is software, it can be installed on the above programmable lamp, and it can be implemented as multiple software or software modules, or it can be implemented as a single software or software module, which is not specifically limited here.

[0079] It is worth mentioning that the user terminal also needs to have a specific video playback function.

[0080] Figure 2 It is the first sub - process block diagram of the data processing method. The steps of real - time obtaining the voice signal sent by the user and performing voice recognition on the voice signal include steps S101 to S104:

[0081] Step S101: Real - time obtain the voice signal sent by the user, input the voice signal into the trained voice recognition model, and obtain the recognition text.

[0082] Step S102: Judge whether the voice signal is an intentional signal according to the recognition text.

[0083] Step S103: When the voice signal is an intentional signal, determine the voice signal information points according to the preset amplitude threshold.

[0084] Step S104: Judge whether the intentional signal is a human voice signal according to the voice signal information points.

[0085] The functions of steps S101 to S104 are to recognize the voice signal sent by the user. Sending a voice signal actually means generating sound, but there are many situations where sound is made, such as the thunder in a thunderstorm. Therefore, before performing voice recognition, it is necessary to judge whether the voice signal is a human voice signal.

[0086] Figure 3 It is the second sub - process block diagram of the data processing method. The steps of judging whether the intentional signal is a human voice signal according to the voice signal information points include steps S1041 to S1044:

[0087] Step S1041: Obtain the amplitudes of a number of voice signal information points;

[0088] Step S1042: Obtain the phase differences of a number of adjacent voice signal information points;

[0089] Step S1043: Calculate the means of the obtained amplitudes and phase differences respectively, and calculate the variance according to the means;

[0090] Step S1044: Compare the calculated variance with a preset variance threshold, and judge whether the intentional signal is a human voice signal according to the comparison result.

[0091] In the technical solution of the present invention, when the user watches a video, a series of sounds will naturally be generated by the video, and these sounds can surely be captured. Therefore, in the process of human voice recognition, an important function is to eliminate these audio signals that are extremely similar to the human voice signal.

[0092] For these recorded audio signals, some feature points are needed for judgment; it can be imagined that the stability of these audio signals is relatively high, while the audio signals actually emitted by the user are often very abrupt. From the perspective of audio signals, it is the difference in the stability of the audio.

[0093] Figure 4 It is the third sub-process block diagram of the data processing method. The steps of obtaining the video file based on the playback request input by the user, performing repetitive recognition on the video file, and determining the adjusted image group according to the repetitive recognition result include Step S301 to Step S304:

[0094] Step S301: Receive the playback request containing the file index input by the user, and extract the video file from the preset file library based on the file index; when the video file does not exist in the preset file library, open the file acquisition port and obtain the video file based on the file acquisition port;

[0095] Step S302: Extract the image information in the video file to generate an image sequence; the number of the image sequence has a mapping relationship with the time information in the image information;

[0096] Step S303: Convert the image sequence into a single-value image library according to a preset conversion formula;

[0097] Step S304: Perform repetitive recognition on the single-value image library, mark the repeated images in the single-value image library based on the repetitive recognition result, and eliminate the repeated images in the video file based on the mapping relationship to obtain the adjusted image group.

[0098] Steps S301 to S304 specifically define the process of identifying duplicate images. First, the acquisition of the image file can be an existing one in the system or uploaded by the user. For the processing of the image file, first, the image file needs to be converted into an image sequence, and the numbering of the image sequence is related to the time corresponding to the image information. Then, perform a color value conversion on the image sequence to obtain a single-value image library. In an example of the technical solution of the present invention, this conversion process can be to convert all the images in the image sequence into grayscale images. Each pixel point in the grayscale image has only one value, and the subsequent comparison process is extremely easy.

[0099] Finally, just remove the duplicate images from the single-value image library.

[0100] Figure 5 It is the fourth sub-process block diagram of the data processing method. The steps of performing duplicate identification on the single-value image library and marking the duplicate images in the single-value image library based on the duplicate identification result include steps S3041 to S3044:

[0101] Step S3041: Read the single-value images in the single-value image library in sequence, traverse each pixel point in the single-value image, and obtain the values of each pixel point.

[0102] Step S3042: Calculate the feature values of the single-value image based on the values of each pixel point to obtain a feature value group corresponding to the single-value image library.

[0103] Step S3043: Calculate the offset rate of adjacent feature values in the feature value group in sequence, and compare the offset rate with a preset offset threshold.

[0104] Step S3044: When the offset rate is less than the preset offset threshold, mark the single-value image corresponding to the subsequent feature value as a duplicate image.

[0105] Steps S3041 to S3044 specifically describe the process of removing duplicate images from the single-value image library. The core of the above content is to calculate the offset rate. If the offset rate of two adjacent images is high, then they are different. If their offset rate is low, it means they are similar in terms of color.

[0106] Figure 6 It is the fifth sub-process block diagram of the data processing method. The steps of inputting the adjusted image group into the lighting adjustment model to generate a lighting adjustment instruction include steps S401 to S403:

[0107] Step S401: Read the feature value group corresponding to the adjusted image group.

[0108] Step S402: Generate a fitting curve based on the eigenvalue group, input the fitting curve into the trained instruction generation model, and obtain a pre-adjustment instruction;

[0109] Step S403: Obtain the adjustment image corresponding to the eigenvalue, obtain the time information of the adjustment image, and insert the time information into the pre-adjustment instruction to obtain a lighting adjustment instruction.

[0110] For the generation process of the lighting adjustment instruction, first, it is necessary to read the eigenvalue group generated during the repeated image rejection process, and then generate a pre-adjustment instruction based on the eigenvalue group; during the process of generating the pre-adjustment instruction, it is necessary to convert different eigenvalues and their time information into a fitting curve, which can make the final brightness adjustment process smooth.

[0111] As a preferred embodiment of the technical solution of the present invention, the method further includes:

[0112] Receive the brightness setting request sent by the user and open the brightness setting port;

[0113] Receive the touch screen signal input by the user based on the brightness setting port, and determine the brightness setting value based on the touch screen signal;

[0114] Replace the reference brightness according to the brightness setting value.

[0115] The key point of the technical solution of the present invention is to finely adjust the ambient brightness in real time according to the image file on the preset reference brightness. Among them, the adjustable range of the reference brightness is much larger than the fine adjustment range. The adjustment of the reference brightness is determined by the user's voice signal, and the granularity of the adjustment process is very large and the adjustment difficulty is also very large; therefore, the above content provides a reference brightness adjustment process with a smaller granularity.

[0116] Embodiment 2

[0117] Figure 7 It is a block diagram of the composition structure of a data processing system. In the embodiment of the present invention, a data processing system, the system 10 includes:

[0118] A voice recognition module 11, configured to obtain the voice signal sent by the user in real time and perform voice recognition on the voice signal; when the voice signal is a voice signal, convert the voice signal into an audio signal;

[0119] A reference adjustment module 12, configured to perform content recognition on the audio signal, determine the reference brightness and the corresponding lighting adjustment model, generate a reference adjustment instruction based on the reference brightness, and send the reference adjustment instruction to the execution end;

[0120] The repeated image recognition module 13 obtains an image file based on a playback request input by a user, performs repeated recognition on the image file, and determines an adjusted image group according to the result of the repeated recognition.

[0121] The instruction generation module 14 is configured to input the adjusted image group into the lighting adjustment model, generate a lighting adjustment instruction, and send the lighting adjustment instruction to an execution end.

[0122] Figure 8 It is a structural block diagram of a voice recognition module in a data processing system. The voice recognition module 11 includes:

[0123] The text generation unit 111 is configured to obtain a voice signal sent by a user in real time, input the voice signal into a trained voice recognition model, and obtain a recognized text.

[0124] The meaning determination unit 112 is configured to determine whether the voice signal is a meaningful signal according to the recognized text.

[0125] The information point determination unit 113 is configured to, when the voice signal is a meaningful signal, determine voice signal information points according to a preset amplitude threshold.

[0126] The processing execution unit 114 is configured to determine whether the meaningful signal is a human voice signal according to the voice signal information points.

[0127] Figure 9 It is a structural block diagram of a processing execution unit in a voice recognition module. The processing execution unit 114 includes:

[0128] The amplitude acquisition sub-unit 1141 is configured to acquire amplitudes of a plurality of voice signal information points.

[0129] The phase difference acquisition sub-unit 1142 is configured to acquire phase differences between a plurality of adjacent voice signal information points.

[0130] The variance calculation sub-unit 1143 is configured to calculate means of the acquired amplitudes and the phase differences respectively, and calculate a variance according to the means.

[0131] The comparison sub-unit 1144 is configured to compare the calculated variance with a preset variance threshold, and determine whether the meaningful signal is a human voice signal according to the comparison result.

[0132] All functions that the data processing method can implement are completed by a computer device. The computer device includes one or more processors and one or more memories. At least one program code is stored in the one or more memories, and the program code is loaded and executed by the one or more processors to implement the functions of the data processing method.

[0133] The processor fetches instructions from the memory one by one, analyzes the instructions, and then completes corresponding operations according to the requirements of the instructions, generating a series of control commands to make each part of the computer act automatically, continuously and coordinately, becoming an organic whole, realizing the input of the program, the input of data, as well as the operation and output of results. All arithmetic operations or logical operations generated in this process are completed by the arithmetic unit; the memory includes a read-only memory (ROM), and the read-only memory is used to store computer programs, and a protection device is provided outside the memory.

[0134] Exemplarily, the computer program can be divided into one or more modules, and one or more modules are stored in the memory and executed by the processor to complete the present invention. One or more modules can be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.

[0135] Those skilled in the art can understand that the description of the above service device is only an example and does not constitute a limitation on the terminal device. It may include more or fewer components than the above description, or combine some components, or different components. For example, it may include input / output devices, network access devices, buses, etc.

[0136] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The above processor is the control center of the above terminal device, and connects various parts of the entire user terminal through various interfaces and lines.

[0137] The above-mentioned memory can be used to store computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory and invoking the data stored in the memory, the above-mentioned processor realizes various functions of the terminal device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as information collection template display function, product information publishing function, etc.); the data storage area can store data created according to the use of the berth status display system (such as product information collection templates corresponding to different product types, product information that different product providers need to publish, etc.). In addition, the memory can include high-speed random access memory and can also include non-volatile memory, such as hard disks, memory, plug-in hard disks, smart media cards (SMC), secure digital (SD) cards, flash cards, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0138] If the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the modules / units in the above-mentioned embodiment system of the present invention, it can also be completed by instructing relevant hardware through a computer program. The above-mentioned computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the functions of the above-mentioned various system embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0139] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element.

[0140] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A data processing method, characterized in that, The method includes: Obtaining a voice signal sent by a user in real time and performing voice recognition on the voice signal; when the voice signal is a human voice signal, converting the voice signal into an audio signal; Performing content recognition on the audio signal, determining a reference brightness and a corresponding lighting adjustment model, generating a reference adjustment instruction based on the reference brightness, and sending the reference adjustment instruction to an execution end; Obtaining an image file based on a playback request input by a user, performing repeatability recognition on the image file, and determining an adjustment image group according to the repeatability recognition result; Inputting the adjustment image group into the lighting adjustment model, generating a lighting adjustment instruction, and sending the lighting adjustment instruction to an execution end; The step of obtaining an image file based on a playback request input by a user, performing repeatability recognition on the image file, and determining an adjustment image group according to the repeatability recognition result includes: Receiving a playback request containing a file index input by a user, extracting an image file from a preset file library based on the file index; when the image file does not exist in the preset file library, opening a file acquisition port and acquiring an image file based on the file acquisition port; Extracting image information in the image file to generate an image sequence; the number of the image sequence has a mapping relationship with the time information in the image information; Converting the image sequence into a single-value image library according to a preset conversion formula; Performing repeatability recognition on the single-value image library, marking repeated images in the single-value image library based on the repeatability recognition result, and removing the repeated images in the image file based on the mapping relationship to obtain an adjustment image group; The step of performing repeatability recognition on the single-value image library and marking repeated images in the single-value image library based on the repeatability recognition result includes: Sequentially reading single-value images in the single-value image library, traversing each pixel point in the single-value image, and obtaining the value of each pixel point; Calculating a feature value of the single-value image based on the value of each pixel point to obtain a feature value group corresponding to the single-value image library; Sequentially calculating the offset rate of adjacent feature values in the feature value group, and comparing the offset rate with a preset offset threshold; When the offset rate is less than the preset offset threshold, marking the single-value image corresponding to the subsequent feature value as a repeated image.

2. The data processing method according to claim 1, characterized in that, The step of obtaining a voice signal sent by a user in real time and performing voice recognition on the voice signal includes: Obtaining a voice signal sent by a user in real time, inputting the voice signal into a trained voice recognition model, and obtaining a recognition text; Judging whether the voice signal is a deliberate signal according to the recognition text; When the voice signal is a deliberate signal, determining voice signal information points according to a preset amplitude threshold; Judging whether the deliberate signal is a human voice signal according to the voice signal information points.

3. The data processing method according to claim 2, characterized in that, The step of judging whether the deliberate signal is a human voice signal according to the voice signal information points includes: Obtaining the amplitudes of a plurality of voice signal information points; Obtaining the phase differences of a plurality of adjacent voice signal information points; Respectively calculating the means of the obtained amplitudes and the phase differences, and calculating a variance according to the means; Compare the calculated variance with a preset variance threshold, and determine whether the intended signal is a human voice signal according to the comparison result.

4. The data processing method according to claim 1, characterized in that, The step of inputting the adjusted image group into the lighting adjustment model to generate a lighting adjustment instruction includes: Read the eigenvalue group corresponding to the adjusted image group; Generate a fitting curve based on the eigenvalue group, input the fitting curve into the trained instruction generation model to obtain a pre-adjustment instruction; Obtain the adjusted image corresponding to the eigenvalue, obtain the time information of the adjusted image, and insert the time information into the pre-adjustment instruction to obtain a lighting adjustment instruction.

5. The data processing method according to any one of claims 1-4, characterized in that, The method further includes: Receive a brightness setting request sent by the user and open a brightness setting port; Receive a touch screen signal input by the user based on the brightness setting port, and determine a brightness setting value based on the touch screen signal; Replace the reference brightness according to the brightness setting value.

6. A data processing system for implementing the method according to any one of claims 1-5, characterized in that, The system includes: A voice recognition module, configured to obtain a voice signal sent by the user in real time and perform human voice recognition on the voice signal; when the voice signal is a human voice signal, convert the voice signal into an audio signal; A reference adjustment module, configured to perform content recognition on the audio signal, determine a reference brightness and a corresponding lighting adjustment model, generate a reference adjustment instruction based on the reference brightness, and send the reference adjustment instruction to the execution end; A repeated image recognition module, configured to obtain an image file based on a play request input by the user, perform repeated recognition on the image file, and determine an adjusted image group according to the repeated recognition result; An instruction generation module, configured to input the adjusted image group into the lighting adjustment model to generate a lighting adjustment instruction, and send the lighting adjustment instruction to the execution end.

7. The data processing system according to claim 6, characterized in that, The voice recognition module includes: A text generation unit, configured to obtain a voice signal sent by the user in real time, input the voice signal into a trained voice recognition model to obtain a recognition text; A meaning determination unit, configured to determine whether the voice signal is an intended signal according to the recognition text; An information point determination unit, configured to determine voice signal information points according to a preset amplitude threshold when the voice signal is an intended signal; A processing execution unit, configured to determine whether the intended signal is a human voice signal according to the voice signal information points.

8. The data processing system according to claim 7, characterized in that, The processing execution unit includes: An amplitude acquisition subunit, configured to acquire the amplitudes of several voice signal information points; A phase difference acquisition subunit, configured to acquire the phase differences of several adjacent voice signal information points; A variance calculation subunit, configured to calculate the means of the acquired amplitudes and phase differences respectively, and calculate the variance according to the means; A comparison subunit, configured to compare the calculated variance with a preset variance threshold, and determine whether the intended signal is a human voice signal according to the comparison result.

Citation Information

Patent Citations

  • Multimedia control method, device and terminal

    CN109903783A

  • Multimedia playing method, device and system, and equipment

    CN112040290A