Image decoding method and device based on reinforcement learning, and electronic equipment
By using a reinforcement learning-based image decoding method and the Q-learning algorithm to dynamically update image parameter adjustment strategies, the problem of low decoding efficiency of traditional methods under complex lighting conditions is solved, and the real-time performance and reliability of image decoding are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING TSINGTENG MICROSYSTEM CO LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional static exposure and gain parameter setting methods are difficult to adapt to rapidly changing lighting conditions, resulting in images that are either too dark or too bright, which greatly reduces decoding efficiency and success rate, especially in scenes with drastic changes in lighting.
An image decoding method based on reinforcement learning is adopted, which dynamically updates the image parameter adjustment strategy through the Q-learning algorithm. The exposure and gain parameters are adjusted according to the image brightness calculation and decoding results to achieve closed-loop adaptive adjustment.
It significantly improves the real-time performance and reliability of image decoding, increases the decoding success rate under complex lighting conditions, and provides more stable and efficient image decoding support.
Smart Images

Figure CN121967708A_ABST
Abstract
Description
Image decoding methods, devices, and electronic equipment based on reinforcement learning Technical Field
[0001] This application relates to the field of image processing technology, such as an image decoding method and apparatus based on reinforcement learning, and electronic equipment. Background Technology
[0002] Currently, with the widespread adoption of IoT, automation, and AI applications, embedded devices extensively utilize image decoding technology to analyze visual information acquired from various sensors or cameras. The accuracy of decoding is crucial to the functionality of embedded systems. However, in application scenarios with complex or frequently changing lighting conditions, traditional static exposure and gain parameter setting methods struggle to adapt to diverse lighting environments, directly impacting image quality and the reliability of subsequent decoding.
[0003] The relevant technologies employ manual parameter adjustment or pre-set fixed parameters to cope with application scenarios with complex or frequently changing lighting conditions.
[0004] In implementing the embodiments of this disclosure, it was found that the related technology has at least the following problems: traditional static parameter settings are difficult to respond to rapidly changing lighting conditions, resulting in images that are either too dark or too bright, which greatly reduces decoding efficiency and success rate. This is especially evident in scenarios involving switching between indoor and outdoor environments or drastic changes in lighting (such as from strong light to weak light, or from shadow to sunlight).
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.
[0007] This disclosure provides an image decoding method, apparatus, and electronic device based on reinforcement learning to improve the accuracy of image decoding.
[0008] In some embodiments, the reinforcement learning-based image decoding method includes: determining an image parameter adjustment strategy corresponding to a first state based on a first state of a first image; adjusting the first image according to the image parameter adjustment strategy and decoding the adjusted first image; and updating the image parameter adjustment strategy using a Q-learning algorithm based on the decoding result of the first image.
[0009] Optionally, based on the first state of the acquired first image, determining an image parameter adjustment strategy corresponding to the first state includes: performing brightness calculation on the acquired first image to obtain the first state of the first image; and determining an image parameter adjustment strategy corresponding to the first state based on the first state.
[0010] Optionally, performing brightness calculation on the acquired first image to obtain a first state of the first image includes: acquiring the first image through a camera or sensor; calculating the brightness value of the first image; and obtaining the first state corresponding to the brightness value.
[0011] Optionally, adjusting the first image according to an image parameter adjustment strategy and decoding the adjusted first image includes: determining a first action corresponding to the image parameter adjustment strategy; wherein the first action includes adjusting exposure parameters and adjusting gain parameters; adjusting the exposure parameters and gain parameters of the first image according to the first action; and decoding the adjusted first image.
[0012] Optionally, based on the decoding result of the first image, the image parameter adjustment strategy is updated using the Q-learning algorithm, including: obtaining a reward value corresponding to the decoding result of the first image; obtaining the second state of the acquired second image and determining the maximum weight value corresponding to the second state; updating the weight value of the Q-learning algorithm using the weight value update formula based on the reward value and the maximum weight value, so as to update the image parameter adjustment strategy.
[0013] Optionally, based on the decoding result of the first image, a reward value corresponding to the decoding result is obtained, including: if the decoding is successful, the reward value is positive; if the decoding fails, the reward value is negative.
[0014] Optionally, the image parameter adjustment strategy is updated, including: adjusting the exposure and gain parameters of the image parameter adjustment strategy, and decoding the newly acquired image after the adjustment is completed.
[0015] In some embodiments, the reinforcement learning-based image decoding device includes: an image acquisition module configured to determine an image parameter adjustment strategy corresponding to a first state based on a first state of the acquired first image; an image decoding module configured to adjust the first image according to the image parameter adjustment strategy and decode the adjusted first image; and a strategy update module configured to update the image parameter adjustment strategy using a Q-learning algorithm based on the decoding result of the first image.
[0016] In some embodiments, the reinforcement learning-based image decoding apparatus includes a processor and a memory storing program instructions, the processor being configured to execute the reinforcement learning-based image decoding method as described above when the program instructions are executed.
[0017] In some embodiments, the electronic device includes: an electronic device body; and an image decoding device based on reinforcement learning, as described above, mounted on the electronic device body.
[0018] The image decoding method, apparatus, and electronic device based on reinforcement learning provided in this disclosure can achieve the following technical effects: In this disclosure, firstly, based on the first state of the acquired first image, a matching image parameter adjustment strategy is accurately located; then, the parameters of the first image are adjusted according to the determined image parameter adjustment strategy, and a decoding operation is performed on the adjusted image. Targeted adjustments ensure that the image parameters are adapted to the current scene, providing a high-quality image foundation for subsequent decoding; finally, combined with the decoding result of the first image, the Q-learning algorithm is used to dynamically update the image parameter adjustment strategy, enabling the strategy to be continuously optimized based on the actual decoding effect. This disclosure achieves closed-loop adaptive image parameter adjustment, significantly improving the real-time performance and reliability of decoding in embedded image decoding scenarios with varying lighting and complex scenes. Simultaneously, through the iterative update characteristics of reinforcement learning, the system can adapt to different environments over a long period, gradually improving the decoding success rate and providing more stable and efficient technical support for the image decoding function of embedded devices.
[0019] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description
[0020] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a limitation on scale, and wherein: FIG1 is a schematic diagram of an image decoding method based on reinforcement learning provided by an embodiment of the present disclosure; FIG2 is a schematic diagram of another image decoding method based on reinforcement learning provided by an embodiment of the present disclosure; FIG3 is a schematic diagram of another image decoding method based on reinforcement learning provided by an embodiment of the present disclosure; FIG4 is a schematic diagram of an image decoding apparatus based on reinforcement learning provided by an embodiment of the present disclosure; FIG5 is a schematic diagram of another image decoding apparatus based on reinforcement learning provided by an embodiment of the present disclosure. Detailed Implementation
[0021] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.
[0022] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.
[0023] Unless otherwise stated, the term "multiple" means two or more.
[0024] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.
[0025] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0026] The term "correspondence" can refer to an association or binding relationship. The correspondence between A and B means that there is an association or binding relationship between A and B.
[0027] With the widespread adoption of IoT, automation, and AI applications, embedded devices extensively utilize image decoding technology to analyze visual information acquired from various sensors or cameras. The accuracy of decoding is crucial to the functionality of embedded systems. However, in application scenarios with complex or frequently changing lighting conditions, traditional static exposure and gain parameter setting methods struggle to adapt to diverse lighting environments, directly impacting image quality and the reliability of subsequent decoding.
[0028] In existing technologies, static parameter settings are difficult to respond to rapidly changing lighting conditions, resulting in images that are either too dark or too bright, which greatly reduces decoding efficiency and success rate. This is especially noticeable in scenarios involving switching between indoor and outdoor environments or drastic changes in lighting (such as from strong light to weak light, or from shadow to sunlight).
[0029] Under complex lighting conditions, if exposure and gain parameters deviate from ideal values, image quality will drop significantly, leading to noticeable fluctuations in decoding success rate and severely impacting the reliability and stability of the equipment in a production environment. Therefore, a highly reliable automatic adjustment technology is needed to significantly improve decoding stability and accuracy.
[0030] Referring to Figure 1, this embodiment of the present disclosure provides an image decoding method based on reinforcement learning, including: S101, determining an image parameter adjustment strategy corresponding to the first state based on the first state of the acquired first image.
[0031] S102, adjust the first image according to the image parameter adjustment strategy, and decode the adjusted first image.
[0032] S103, based on the decoding result of the first image, update the image parameter adjustment strategy using the Q-learning algorithm.
[0033] The method provided in this disclosure uses the first state of the acquired first image as a basis to accurately locate a matching image parameter adjustment strategy. Then, the parameters of the first image are adjusted according to the determined image parameter adjustment strategy, and a decoding operation is performed on the adjusted image. Targeted adjustments ensure that the image parameters are adapted to the current scene, providing a high-quality image foundation for subsequent decoding. Finally, based on the decoding result of the first image, the Q-learning algorithm is used to dynamically update the image parameter adjustment strategy, enabling the strategy to be continuously optimized based on the actual decoding effect. This disclosure achieves closed-loop adaptive image parameter adjustment. In embedded image decoding scenarios with varying lighting and complex scenes, it can significantly improve the real-time performance and reliability of decoding. Simultaneously, through the iterative update characteristics of reinforcement learning, the system can adapt to different environments over a long period, gradually improving the decoding success rate and providing more stable and efficient technical support for the image decoding function of embedded devices.
[0034] Optionally, based on the first state of the acquired first image, determining an image parameter adjustment strategy corresponding to the first state includes: performing brightness calculation on the acquired first image to obtain the first state of the first image; and determining an image parameter adjustment strategy corresponding to the first state based on the first state.
[0035] In this embodiment, brightness calculation is first performed on the acquired first image. By analyzing the brightness information of the image pixels, a first state that reflects the image lighting conditions is accurately obtained. This process provides a quantitative basis for subsequent strategy matching. Based on the obtained first state, a corresponding image parameter adjustment strategy is selected from a preset parameter adjustment strategy library to ensure that the strategy is highly consistent with the actual lighting conditions of the image. Through a clear brightness calculation step, the deviation of the adjustment strategy caused by misjudgment of the state is reduced, thereby improving the accuracy of parameter adjustment.
[0036] Optionally, performing brightness calculation on the acquired first image to obtain a first state of the first image includes: acquiring the first image through a camera or sensor; calculating the brightness value of the first image; and obtaining the first state corresponding to the brightness value.
[0037] In this embodiment, a first image is acquired using a camera or sensor. An image brightness calculation algorithm is then used to quantify and analyze the overall brightness of the first image, obtaining specific brightness values. This transforms the image's illumination into quantifiable numerical values, providing precise data support for state determination. Based on preset brightness value-state correspondence rules, the calculated brightness values are matched against these rules to ultimately determine the first state corresponding to the first image. By clearly defining the image acquisition device type, the method's hardware adaptability is enhanced, making it compatible with commonly used image acquisition modules in embedded systems. The quantification of brightness values avoids the errors of traditional qualitative image brightness judgments, improving the accuracy of state determination. The state matching rules reduce human interference, ensuring consistency in state determination across different scenarios. This allows subsequent parameter adjustment strategies to better align with the actual image conditions, effectively solving the problems of ambiguous brightness determination and inaccurate state division.
[0038] Optionally, adjusting the first image according to an image parameter adjustment strategy and decoding the adjusted first image includes: determining a first action corresponding to the image parameter adjustment strategy; wherein the first action includes adjusting exposure parameters and adjusting gain parameters; adjusting the exposure parameters and gain parameters of the first image according to the first action; and decoding the adjusted first image.
[0039] In this embodiment, firstly, a corresponding first action is extracted from the determined image parameter adjustment strategy. This first action includes adjusting exposure parameters and adjusting gain parameters to ensure comprehensive adjustment. Secondly, according to the specific adjustment requirements for exposure and gain parameters in the first action, these two parameters of the first image are precisely adjusted. Targeted parameter adjustments improve key quality indicators such as image brightness and contrast. Finally, a professional image decoding algorithm is used to perform a decoding operation on the parameter-adjusted first image to verify the effect of the parameter adjustment and the image's decodeability. The parameter adjustment in this embodiment can quickly optimize image quality, making the image more compatible with the input requirements of the decoding algorithm and improving the decoding success rate. Furthermore, the decoding operation provides real-time feedback on the effect of parameter adjustment, providing a basis for subsequent updates to the adjustment strategy using the Q-learning algorithm, thus improving the effectiveness of decoding.
[0040] Optionally, based on the decoding result of the first image, the image parameter adjustment strategy is updated using the Q-learning algorithm, including: obtaining a reward value corresponding to the decoding result of the first image; obtaining the second state of the acquired second image and determining the maximum weight value corresponding to the second state; updating the weight value of the Q-learning algorithm using the weight value update formula based on the reward value and the maximum weight value, so as to update the image parameter adjustment strategy.
[0041] Referring to Figure 2, this embodiment of the present disclosure provides another image decoding method based on reinforcement learning, including: S201, determining an image parameter adjustment strategy corresponding to the first state based on the first state of the acquired first image.
[0042] S202, adjust the first image according to the image parameter adjustment strategy, and decode the adjusted first image.
[0043] S203, based on the decoding result of the first image, obtain the reward value corresponding to the decoding result.
[0044] S204, acquire the second state of the acquired second image, and determine the maximum weight value corresponding to the second state.
[0045] S205, based on the reward value and the maximum weight value, the weight values of the Q-learning algorithm are updated using the weight value update formula to update the image parameter adjustment strategy.
[0046] In this embodiment, based on the decoding result (success or failure) of the first image and combined with a preset reward rule, a corresponding reward value is obtained, converting the decoding effect into a quantized signal recognizable by the reinforcement learning algorithm. A new second image is acquired, and the second state corresponding to the second image is determined using the same state determination method as the first image. From a preset weight value table, all weight values corresponding to the second state are selected, and the maximum weight value is determined, which represents the optimal action selection tendency in the second state. The reward value and the maximum weight value are substituted into the weight value update formula of the Q-learning algorithm to calculate and update the weight value corresponding to the current strategy, thereby completing the iterative optimization of the image parameter adjustment strategy. Quantifying the decoding result through the reward value enables the reinforcement learning algorithm to accurately capture the advantages and disadvantages of the adjustment strategy, ensuring the correctness of the update direction. The dynamic update of the weight value realizes the continuous optimization of the adjustment strategy, enabling the system to gradually accumulate optimal adjustment experience in different scenarios, significantly enhancing the adaptability of the method in complex and ever-changing environments.
[0047] Optionally, based on the decoding result of the first image, a reward value corresponding to the decoding result is obtained, including: if the decoding is successful, the reward value is positive; if the decoding fails, the reward value is negative.
[0048] Referring to Figure 3, this embodiment of the present disclosure provides another image decoding method based on reinforcement learning, including: S301, determining an image parameter adjustment strategy corresponding to the first state based on the first state of the acquired first image.
[0049] S302, adjust the first image according to the image parameter adjustment strategy, and decode the adjusted first image.
[0050] S303, if decoding is successful, the reward value obtained is a positive number.
[0051] S304: In the event of decoding failure, the reward value obtained is negative.
[0052] S305, acquire the second state of the acquired second image, and determine the maximum weight value corresponding to the second state.
[0053] S306, based on the reward value and the maximum weight value, the weight values of the Q-learning algorithm are updated using the weight value update formula to update the image parameter adjustment strategy.
[0054] In this embodiment, when the first image is successfully decoded, the adjustment strategy is determined to have a positive effect on decoding, and a positive reward value is assigned to strengthen the effectiveness of the strategy from a positive incentive perspective. When the first image fails to decode, it indicates that the current adjustment strategy is insufficient or unsuitable for the current scenario, and a negative reward value is assigned to prompt optimization of the strategy from a negative constraint perspective. By dividing the reward into positive and negative values, the reinforcement learning algorithm can quickly and intuitively determine the effect of the current adjustment strategy, reducing the complexity of processing feedback signals and adapting to the limited computing resources of embedded devices. Positive rewards can increase the frequency of using high-quality strategies, while negative rewards can encourage the system to reduce the selection of ineffective strategies, guiding the adjustment strategy to iterate towards improving the decoding success rate. This avoids the problem of lack of clear feedback guidance in strategy optimization in traditional methods, making strategy updates more targeted and effective, further accelerating the system's adaptation to different scenarios, and improving overall decoding efficiency.
[0055] Optionally, the weight value update formula is as follows: .
[0056] in, The weight value corresponding to the action in the current state. For learning rate, As a reward value, As a discount factor, This represents the maximum weight value corresponding to all actions in the next state.
[0057] In this embodiment of the disclosure, the Q-learning algorithm in the field of reinforcement learning is adopted. The weight value corresponding to the action is updated in real time through the weight value update formula, thereby continuously improving the accuracy of the exposure and gain parameter adjustment strategy to achieve efficient and stable image decoding.
[0058] Optionally, the image parameter adjustment strategy is updated, including: adjusting the exposure and gain parameters of the image parameter adjustment strategy, and decoding the newly acquired image after the adjustment is completed.
[0059] In this embodiment, a new weight value is calculated, which directly reflects the adaptation effect of the current combination of exposure and gain parameters. Based on the new weight value, the exposure and gain parameters in the image parameter adjustment strategy are adjusted accordingly. If the new weight value is higher than before the update, it indicates that the current parameter combination has good adaptation, and small optimizations can be made within this parameter range to consolidate the effect; if the new weight value is lower than before the update, the parameter adjustment direction needs to be adjusted according to the magnitude of the weight value change (such as increasing or decreasing the exposure time, increasing or decreasing the gain factor) to form an updated parameter combination, completing the iteration of the image parameter adjustment strategy. Finally, after the exposure and gain parameters are adjusted, a new image is acquired in real time through the image acquisition module, and a decoding operation is performed on the new image to obtain the decoding result after this parameter adjustment. This result will serve as the feedback basis for the next round of weight value updates, forming a complete closed loop of "weight update, parameter adjustment, and decoding verification".
[0060] In this way, by incorporating the reward value and the maximum weight value into the weight update process of the Q-learning algorithm, the adjustment of exposure and gain parameters has clear algorithmic support, avoiding the blindness of parameter adjustment and significantly improving the accuracy of parameter tuning. The dynamic updating of weight values allows image parameter adjustment strategies to solve the problem that traditional fixed-parameter strategies cannot adapt to environmental changes.
[0061] Referring to Figure 4, this embodiment of the present disclosure provides an image decoding device 400 based on reinforcement learning, including an image acquisition module 401, an image decoding module 402, and a policy update module 403. The image acquisition module 401 is configured to determine an image parameter adjustment policy corresponding to a first state of the acquired first image; the image decoding module 402 is configured to adjust the first image according to the image parameter adjustment policy and decode the adjusted first image; the policy update module 403 is configured to update the image parameter adjustment policy using a Q-learning algorithm based on the decoding result of the first image.
[0062] The image decoding device 400 based on reinforcement learning provided in this disclosure firstly, based on the first state of the acquired first image, accurately locates the matching image parameter adjustment strategy; then, it adjusts the parameters of the first image according to the determined image parameter adjustment strategy, and performs a decoding operation on the adjusted image. This targeted adjustment ensures that the image parameters are adapted to the current scene, providing a high-quality image foundation for subsequent decoding; finally, based on the decoding result of the first image, the Q-learning algorithm is used to dynamically update the image parameter adjustment strategy, enabling the strategy to be continuously optimized according to the actual decoding effect. This disclosure achieves closed-loop adaptive image parameter adjustment, significantly improving the real-time performance and reliability of decoding in embedded image decoding scenarios with varying lighting and complex scenes. Simultaneously, through the iterative update characteristics of reinforcement learning, the system can adapt to different environments over a long period, gradually improving the decoding success rate and providing more stable and efficient technical support for the image decoding function of embedded devices.
[0063] Referring to Figure 5, this embodiment of the present disclosure provides an image decoding device 50 based on reinforcement learning, including a processor 500 and a memory 501. Optionally, the device 50 may further include a communication interface 502 and a bus 503. The processor 500, communication interface 502, and memory 501 can communicate with each other via the bus 503. The communication interface 502 can be used for information transmission. The processor 500 can call logical instructions in the memory 501 to execute the reinforcement learning-based image decoding method of the above embodiment.
[0064] Furthermore, the logic instructions in the aforementioned memory 501 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.
[0065] The memory 501, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 500 executes functional applications and data processing by running the program instructions / modules stored in the memory 501, that is, it implements the reinforcement learning-based image decoding method in the above embodiments.
[0066] The memory 501 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 501 may include high-speed random access memory and may also include non-volatile memory.
[0067] This disclosure provides an electronic device, including: an electronic device body, and the aforementioned reinforcement learning-based image decoding device. The reinforcement learning-based image decoding device is mounted on the electronic device body. The mounting relationship described herein is not limited to placement within the electronic device body, but also includes mounting connections with other components of the electronic device, including but not limited to physical connections, electrical connections, or signal transmission connections. Those skilled in the art will understand that the reinforcement learning-based image decoding device can be adapted to feasible electronic device bodies to achieve other feasible embodiments.
[0068] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this disclosure. The aforementioned storage medium can be a non-transitory storage medium, such as a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., and other media capable of storing program code.
[0069] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.
[0070] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0071] The methods and products disclosed in the embodiments herein (including but not limited to devices and equipment) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0072] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
Claims
1. An image decoding method based on reinforcement learning, characterized in that, include: Based on the first state of the acquired first image, determine the image parameter adjustment strategy corresponding to the first state; The first image is adjusted according to the image parameter adjustment strategy, and the adjusted first image is decoded; based on the decoding result of the first image, the image parameter adjustment strategy is updated using the Q-learning algorithm.
2. The image decoding method according to claim 1, characterized in that, Based on the first state of the acquired first image, determine the image parameter adjustment strategy corresponding to the first state, including: performing brightness calculation on the acquired first image to obtain the first state of the first image; and determining the image parameter adjustment strategy corresponding to the first state based on the first state.
3. The image decoding method according to claim 2, characterized in that, The process of calculating the brightness of the first image to obtain a first state of the first image includes: acquiring the first image through a camera or sensor; calculating the brightness value of the first image; and obtaining the first state corresponding to the brightness value.
4. The image decoding method according to claim 1, characterized in that, Adjusting a first image according to an image parameter adjustment strategy and decoding the adjusted first image includes: determining a first action corresponding to the image parameter adjustment strategy; wherein the first action includes adjusting exposure parameters and adjusting gain parameters; adjusting the exposure parameters and gain parameters of the first image according to the first action; and decoding the adjusted first image.
5. The image decoding method according to any one of claims 1 to 4, characterized in that, Based on the decoding result of the first image, the image parameter adjustment strategy is updated using the Q-learning algorithm, including: obtaining the reward value corresponding to the decoding result of the first image; obtaining the second state of the acquired second image and determining the maximum weight value corresponding to the second state; updating the weight value of the Q-learning algorithm using the weight value update formula based on the reward value and the maximum weight value, so as to update the image parameter adjustment strategy.
6. The image decoding method according to claim 5, characterized in that, Based on the decoding result of the first image, obtain the reward value corresponding to the decoding result, including: if the decoding is successful, the reward value is positive; if the decoding fails, the reward value is negative.
7. The image decoding method according to claim 5, characterized in that, The image parameter adjustment strategy is updated, including adjusting the exposure and gain parameters of the image parameter adjustment strategy, and decoding the newly acquired image after the adjustment is completed.
8. An image decoding device based on reinforcement learning, characterized in that, include: The image acquisition module is configured to determine an image parameter adjustment strategy corresponding to the first state based on the first state of the acquired first image. The image decoding module is configured to adjust the first image according to the image parameter adjustment strategy and decode the adjusted first image; The strategy update module is configured to update the image parameter adjustment strategy based on the decoding result of the first image using the Q-learning algorithm.
9. An image decoding device based on reinforcement learning, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute the reinforcement learning-based image decoding method as described in any one of claims 1 to 7 when running the program instructions.
10. An electronic device, characterized in that, include: The electronic device body; the reinforcement learning-based image decoding device as described in claim 8 or 9, is installed on the electronic device body.