Automated Processing Method for Financial Statements Based on Statistical Analysis and Image Recognition

By improving YOLOv8 target recognition network model and statistical analysis technology, the financial statements are automatically identified and analyzed, and the problems of inefficiency and lack of statistical analysis of traditional financial statement processing methods are solved, and efficient and accurate financial data processing and report generation are achieved.

CN119622064BActive Publication Date: 2025-06-24GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411721090.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-06-24
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

Traditional financial statement processing methods are inefficient, prone to errors, high cost, and lack statistical analysis functions for financial data, making it difficult to meet the needs of modern enterprises.

Method used

The automated processing method of financial statements based on statistical analysis and image recognition is adopted to identify key information in financial statements by improving the YOLOv8 target recognition network model, and combine time series analysis and data visualization to generate financial reports.

Benefits of technology

It greatly reduces the time for manually inputting data, improves work efficiency, significantly reduces the error rate, can quickly process a large amount of financial data, and improves the financial management level and competitiveness of the enterprise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119622064B_ABST
    Figure CN119622064B_ABST
Patent Text Reader

Abstract

The present invention relates to an automated processing method for financial statements based on statistical analysis and image recognition, belonging to the field of target image recognition. It includes the following steps: obtaining a financial statement sample image to obtain a financial statement image dataset; constructing a financial statement target recognition network model based on the improved YOLOv8, the model including: an Input input end, a Backbone backbone network, a Neck network, and a Head output end, and the images in the financial statement image dataset are recognized by the model to obtain financial statement image text information; performing pandas data analysis on the financial statement image text information through a time series analysis method, and using the matplotlib and seaborn libraries to process the results of the data analysis to generate relevant trend charts for visualization; based on the results of the data analysis and the generated relevant trend charts, creating a financial report in the Word document format through the python-docx library. The method of the present invention can reduce the time for manually inputting data and improve work efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of target image recognition, and particularly relates to an automated financial statement processing method based on statistical analysis and image recognition. Background Art

[0002] Financial statements are important reflections of a company's financial position and operating results, and their processing and analysis are of great significance for company management and investment decisions. Traditional financial statement processing methods mainly rely on manual operations, which have problems such as low efficiency, easy errors, high costs, and difficulty in achieving standardization, resulting in uneven quality of financial statements. With the increase in enterprise data volume, traditional manual financial statement processing methods can no longer meet the needs of modern enterprises. Manual data entry is prone to errors and slow in processing speed, seriously affecting the timeliness of financial decisions. In recent years, image recognition and statistical analysis technologies have made remarkable progress in multiple fields, yet their applications in automated financial statement processing are still limited.

[0003] Currently, image recognition technology has been widely applied in multiple industries, especially in document processing, automated entry, and intelligent auditing. Through deep learning models such as the YOLO series, a computer can quickly and accurately identify target objects in images. The application of this technology has greatly improved the speed and accuracy of data extraction and reduced manual intervention. However, existing algorithms are limited by the accuracy of image recognition algorithms and are difficult to accurately identify key information in financial statements. Existing technologies often require manual data cleaning and format conversion, with low processing efficiency. Existing technologies mainly focus on information extraction and lack the statistical analysis function for financial data.

[0004] In addition, relying solely on image recognition cannot meet the complete financial data processing requirements, and statistical analysis also needs to be combined for effective interpretation and utilization of data. Statistical analysis methods such as regression analysis, trend analysis, and descriptive statistics have shown important value in data-driven decision-making. However, existing statistical analysis tools often lack the intelligent processing ability for data sources, which makes it impossible for enterprises to obtain accurate financial insights in a rapidly changing market environment. Summary of the Invention

[0005] In order to solve the above problems, the present invention provides an automated financial statement processing method based on statistical analysis and image recognition.

[0006] To achieve the above object, the present invention is realized through the following technical solutions:

[0007] In a first aspect, the present invention provides an automated financial statement processing method based on statistical analysis and image recognition, including the following steps:

[0008] S1. Obtain a sample image of the financial statement to get a financial statement image dataset;

[0009] S2. Construct a financial statement target recognition network model based on the improved YOLOv8. The financial statement target recognition network model based on the improved YOLOv8 includes: an Input input end, a Backbone backbone network, a Neck network, and a Head output end. The images in the financial statement image dataset are recognized by the financial statement target recognition network model based on the improved YOLOv8 to obtain the financial statement image text information;

[0010] S3. Perform pandas data analysis on the financial statement image text information through time series analysis methods to obtain the data analysis results, and generate relevant trend charts from the data analysis results by using the matplotlib and seaborn libraries to achieve visualization;

[0011] S4. Based on the data analysis results and the generated relevant trend charts, create a financial report in the Word document format through the python-docx library.

[0012] Further, step S1 specifically includes:

[0013] Set up an automated script to regularly obtain public financial statement sample data based on the automated script and remove redundant information in the data to obtain a financial statement image dataset .

[0014] Further, the operation of the Input input end in step S2 is specifically:

[0015] The Input input end includes a Mosaic data augmentation module, an adaptive anchor box module, and an adaptive image scaling module;

[0016] Financial statement image dataset The images in it pass through the Input input end to obtain a preprocessed financial statement image dataset , , Denote the th preprocessed financial statement image sample.

[0017] Further, the operation of the Backbone backbone network in step S2 is specifically:

[0018] The Backbone backbone network includes a convolutional Conv module, a C2 feature extraction layer C2f_1, a feature extraction layer C2f_Star module, and an AIFI module; the convolutional Conv module includes Conv_1, Conv_2, Conv_3, Conv_4, Conv_5, Conv6, and the feature extraction layer C2f_Star module includes C2f_Star_1, C2f_Star_2, and C2f_Star_3; C2f_Star_1 includes 1 convolutional Conv1 module, 2 Star Blocks modules, and 1 convolutional Conv2 module; the Star Blocks module includes 1 depthwise separable convolution DW-Conv1, 3 1×1 fully connected layers, and 1 depthwise separable convolution DW-Conv2, and the structures of C2f_Star_2 and C2f_Star_3 are the same as that of C2f_Star_1;

[0019] Image samples in the preprocessed financial statement image dataset The first feature map with 64 output channels is output after passing through the convolutional layer Conv_1 , the first feature map The second feature map with 128 output channels is output after passing through the convolutional layer Conv_2 , the second feature map The third feature map is obtained after passing through the feature extraction layer C2f_1 , the third feature map The fourth feature map with 256 output channels is output after passing through the convolutional layer Conv_3 , the fourth feature map The fifth feature map with 256 output channels is output after passing through the feature extraction layer C2f_Star_1 , the fifth feature map The sixth feature map with 512 output channels is output after passing through the convolutional layer Conv_4 , the sixth feature map The seventh feature map with 512 output channels is obtained after passing through the feature extraction layer C2f_Star_2 , the seventh feature map The eighth feature map with 1024 output channels is output after passing through the convolutional layer Conv_5 , the eighth feature map After passing through the feature extraction layer C2f_Star_3, the ninth feature map with 1024 output channels is obtained , the ninth feature map After downsampling through the convolutional layer Conv6, the number of channels becomes 512, and then it is input into the AIFI module for scale-internal interaction of high-level semantic features to obtain the tenth feature map with 512 output channels .

[0020] Further, the operation of the Neck network in step S2 is specifically as follows:

[0021] The Neck network includes a Feature Focusing Diffusion Pyramid Network (FCF) module, a Convolutional layer (Conv) module, a C2 feature extraction layer, a Concat connection layer, and an Upsampling layer (Upsample) module; the Feature Focusing Diffusion Pyramid Network (FCF) module includes FCF_1 and FCF_2, the Convolutional layer (Conv) module includes Conv_6, Conv_7, and Conv_8, the C2 feature extraction layer includes C2f_2, C2f_3, C2f_4, and C2f_5, the Concat connection layer includes Concat_1, Concat_2, Concat_3, and Concat_4, and the Upsampling layer (Upsample) module includes Upsample_1 and Upsample_2;

[0022] The seventh feature map , the tenth feature map and the fifth feature map pass through the Feature Focusing Diffusion Pyramid Network FCF_1 to obtain an eleventh feature map with an output channel number of 128 , and the eleventh feature map passes through the Convolutional layer Conv_6 to obtain a twelfth feature map with an output channel number of 256 , and the twelfth feature map is connected to the tenth feature map through the Concat connection layer Concat_1 to obtain a thirteenth feature map with an output channel number of 256 , and the thirteenth feature map passes through the feature extraction layer C2f_2 to obtain a fourteenth feature map with an output channel number of 256 , and the eleventh feature map passes through the Upsampling layer Upsample module Upsample_1, is upsampled by a factor of 2, and outputs a fifteenth feature map with an output channel number of 256 , and the fifteenth feature map is concatenated with the fifth feature map through the Concat connection layer Concat_2 to obtain a sixteenth feature map with an output channel number of 256 , and the sixteenth feature map passes through the feature extraction layer C2f_3 to obtain a seventeenth feature map with an output channel number of 256 , and the seventeenth feature map , the eleventh feature map and the fourteenth feature map After passing through the Feature Focused Diffusion Pyramid Network FCF_2, the eighteenth feature map with 256 output channels is obtained , and the eighteenth feature map passes through the Conv_7 module to obtain the nineteenth feature map with 512 output channels , and the nineteenth feature map is concatenated with the twelfth feature map and the fourteenth feature map through the Concat_3 module to obtain the twentieth feature map with 1024 output channels , and the twentieth feature map passes through the feature extraction layer C2f_4 to obtain the twenty-first feature map with 1024 output channels , and the twenty-first feature map passes through the upsampling layer Upsample module Upsample_2 to obtain the twenty-second feature map , and the fifteenth feature map , the twenty-second feature map and the seventeenth feature map are concatenated through the Concat connection layer Concat_3 to obtain the twenty-third feature map with 1024 output channels , and the twenty-third feature map passes through the feature extraction layer C2f_5 to obtain the twenty-fourth feature map with 1024 output channels .

[0023] Furthermore, the operation at the output end of Head in step S2 is specifically as follows:

[0024] The output end of Head consists of the Detect_DyHead module; the Detect_DyHead module includes a convolutional layer Conv module and a DyHeadBlock module;

[0025] The eighteenth feature map and the twenty-fourth feature map pass through the Detect_Dyhead detection module to generate detection frames, complete the recognition of the financial statement image, and obtain the financial statement image text information;

[0026] The financial statement image text information includes financial indicators, timestamps, market data, and macroeconomic data.

[0027] Furthermore, in step S3, it is specifically as follows:

[0028] The data analysis results include the changing trends of financial indicators over time, periodic fluctuations in the data, the correlations between different financial indicators, and the correlations between macroeconomic indicators; the relevant trend charts include monthly revenue comparison charts, monthly target completion rate line charts, monthly completion rate percentage charts, monthly revenue comparison tables, current ratio tables, debt ratio tables, and profit margin tables.

[0029] In a second aspect, a financial statement automated processing device based on statistical analysis and image recognition includes:

[0030] A data acquisition unit: used to acquire financial statement sample images to obtain a financial statement image data set;

[0031] A model construction unit: used to construct a financial statement target recognition network model based on improved YOLOv8. The financial statement target recognition network model based on improved YOLOv8 includes: an Input input end, a Backbone main network, a Neck network, and a Head output end. Images in the financial statement image data set are recognized by the financial statement target recognition network model based on improved YOLOv8 to obtain financial statement image text information;

[0032] A data analysis unit: used to perform pandas data analysis on the financial statement image text information through time series analysis methods to obtain data analysis results, and generate relevant trend charts by using the matplotlib and seaborn libraries to achieve visualization;

[0033] A report generation unit: used to create a financial report in Word document format based on the data analysis results and the generated relevant trend charts through the python-docx library.

[0034] In a third aspect, an electronic device includes: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the above-mentioned financial statement automated processing method based on statistical analysis and image recognition is executed.

[0035] In a fourth aspect, a computer-readable storage medium stores a computer program. When the computer program is run by a processor, the above-mentioned financial statement automated processing method based on statistical analysis and image recognition is executed.

[0036] The advantages of the present invention are:

[0037] The method of the present invention greatly reduces the time for manually inputting data, improving work efficiency. Moreover, statistical analysis and image recognition technologies can accurately identify and input data, significantly reducing the error rate and enabling rapid processing of a large amount of financial data. At the same time, it reduces labor costs, which is a huge advantage for enterprises with a large number of financial statements. Fast, accurate, and efficient financial statement processing can enhance the financial management level of enterprises and strengthen their overall competitiveness. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention.

[0039] Figure 1 is a flowchart of the method of the present invention;

[0040] Figure 2 is the recognition accuracy of different algorithms of the present invention;

[0041] Figure 3 is a chart generated after data analysis of the present invention;

[0042] Figure 4 is the financial statement generated by the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] Embodiment 1

[0045] In this embodiment, as Figure 1 shown, the present invention provides an automated financial statement processing method based on statistical analysis and image recognition. The specific steps include:

[0046] S1. Obtain a financial statement sample image to obtain a financial statement image data set;

[0047] Specifically, set an automated script to regularly obtain public financial statement sample data of Chinese listed companies based on the automated script, and clear the redundant information in the data to obtain a financial statement image data set .

[0048] S2. Construct a financial statement target recognition network model based on the improved YOLOv8. The financial statement target recognition network model based on the improved YOLOv8 includes: an Input input end, a Backbone backbone network, a Neck network, and a Head output end. The images in the financial statement image dataset are recognized by the financial statement target recognition network model based on the improved YOLOv8 to obtain the financial statement image text information;

[0049] Specifically, the Input input end includes a Mosaic data augmentation module, an adaptive anchor box module, and an adaptive image scaling module; the financial statement image dataset The images in it are processed by the Input input end to obtain a preprocessed financial statement image dataset , , denotes the th preprocessed financial statement image sample.

[0050] Specifically, the Backbone backbone network includes a convolutional Conv module, a C2 feature extraction layer C2f_1, a feature extraction layer C2f_Star module, and an AIFI module; the convolutional Conv module includes Conv_1, Conv_2, Conv_3, Conv_4, Conv_5, Conv6, and the feature extraction layer C2f_Star module includes C2f_Star_1, C2f_Star_2, and C2f_Star_3; C2f_Star_1 includes 1 convolutional Conv1 module, 2 Star Blocks modules, and 1 convolutional Conv2 module; the Star Blocks module includes 1 depthwise separable convolution DW-Conv1, 3 1×1 fully connected layers, and 1 depthwise separable convolution DW-Conv2. The structures of C2f_Star_2 and C2f_Star_3 are the same as that of C2f_Star_1;

[0051] The image samples in the preprocessed financial statement image dataset pass through the convolutional layer Conv_1 to output the first feature map with 64 output channels , and the first feature map passes through the convolutional layer Conv_2 to output the second feature map with 128 output channels , and the second feature map passes through the feature extraction layer C2f_1 to obtain the third feature map , and the third feature map passes through the convolutional layer Conv_3 to output the fourth feature map with 256 output channels , and the fourth feature map The fifth feature map with 256 output channels is output by the feature extraction layer C2f_Star_1 , the fifth feature map The sixth feature map with 512 output channels is output by the convolutional layer Conv_4 , the sixth feature map The seventh feature map with 512 output channels is obtained by the feature extraction layer C2f_Star_2 , the seventh feature map The eighth feature map with 1024 output channels is output by the convolutional layer Conv_5 , the eighth feature map The ninth feature map with 1024 output channels is obtained by the feature extraction layer C2f_Star_3 , the ninth feature map After downsampling through the convolutional layer Conv6, the number of channels becomes 512, and then it is input into the AIFI module for scale-internal interaction of high-level semantic features, obtaining the tenth feature map with 512 output channels .

[0052] Specifically, the Neck network includes a Feature Focusing and Diffusion Pyramid Network FCF module, a convolutional layer Conv module, a C2 feature extraction layer, a Concat connection layer, and an upsampling layer Upsample module; the Feature Focusing and Diffusion Pyramid Network FCF module includes FCF_1 and FCF_2, the convolutional layer Conv module includes Conv_6, Conv_7, and Conv_8, the C2 feature extraction layer includes C2f_2, C2f_3, C2f_4, and C2f_5, the Concat connection layer includes Concat_1, Concat_2, Concat_3, and Concat_4, and the upsampling layer Upsample module includes Upsample_1 and Upsample_2;

[0053] The seventh feature map , the tenth feature map and the fifth feature map Pass through the Feature Focusing and Diffusion Pyramid Network FCF_1 to obtain the eleventh feature map with 128 output channels , the eleventh feature map Pass through the convolutional layer Conv_6 to obtain the twelfth feature map with 256 output channels , the twelfth feature map and the tenth feature map Are connected through the Concat connection layer Concat_1 to obtain the thirteenth feature map with 256 output channels , the thirteenth feature map After passing through the feature extraction layer C2f_2, the fourteenth feature map with 256 output channels is obtained , the eleventh feature map passes through the upsampling layer Upsample module Upsample_1, is upsampled by a factor of 2, and the output is the fifteenth feature map with 256 output channels , the fifteenth feature map and the fifth feature map are concatenated through the Concat connection layer Concat_2 to obtain the sixteenth feature map with 256 output channels , the sixteenth feature map passes through the feature extraction layer C2f_3 to obtain the seventeenth feature map with 256 output channels , the seventeenth feature map , the eleventh feature map and the fourteenth feature map pass through the Feature Focusing Diffusion Pyramid Network FCF_2 to obtain the eighteenth feature map with 256 output channels , the eighteenth feature map passes through the Conv_7 module to obtain the nineteenth feature map with 512 output channels , the nineteenth feature map and the twelfth feature map , the fourteenth feature map are concatenated through the Concat_3 module to obtain the twentieth feature map with 1024 output channels , the twentieth feature map passes through the feature extraction layer C2f_4 to obtain the twenty-first feature map with 1024 output channels , the twenty-first feature map passes through the upsampling layer Upsample module Upsample_2 to obtain the twenty-second feature map , the fifteenth feature map , the twenty-second feature map and the seventeenth feature map are concatenated through the Concat connection layer Concat_3 to obtain the twenty-third feature map with 1024 output channels , the twenty-third feature map passes through the feature extraction layer C2f_5 to obtain the twenty-fourth feature map with 1024 output channels .

[0054] Specifically, the output end of the Head consists of the Detect_DyHead module; the Detect_DyHead module includes a convolutional layer Conv module and a DyHeadBlock module;

[0055] The eighteenth feature map and the twenty-fourth feature map are detected by the Detect_Dyhead detection module to generate detection frames, complete the recognition of the financial statement image, and obtain the text information of the financial statement image;

[0056] The text information of the financial statement image includes financial indicators, timestamps, market data, and macroeconomic data.

[0057] S3. Perform pandas data analysis on the text information of the financial statement image through time series analysis methods to obtain the data analysis results, and generate relevant trend charts from the data analysis results by using the matplotlib and seaborn libraries to achieve visualization;

[0058] Specifically, the data analysis results include the change trend of financial indicators over time, periodic fluctuations in the data, the correlation between different financial indicators, and the correlation between macroeconomic indicators; the relevant trend charts include monthly income comparison charts, monthly target completion rate line charts, monthly completion rate percentage charts, monthly income comparison tables, current ratio tables, debt ratio tables, and profit margin tables.

[0059] S4. Based on the data analysis results and the generated relevant trend charts, create a financial report in Word document format through the python-docx library.

[0060] Embodiment 2

[0061] In this embodiment, Table 1 shows the experimental comparison data between the financial statement target recognition network model based on the improved YOLOv8 of the present invention and other detection algorithms. It can be seen from Table 1 that the recognition rate and accuracy of the proposed model are higher than those of the existing models, which can better improve work efficiency, meet the financial management needs of various enterprises, and have wide practicability and promotion value.

[0062] Table 1 Experimental comparison data between the method of the present invention and other detection algorithms

[0063]

[0064] In this embodiment, Figure 2 shows the comparison of the recognition accuracy of the method of the present invention and other detection algorithms at a time point. It can be seen from Figure 2 that the improved model has the highest recognition accuracy, which is higher than that of the existing models, is fast, can obtain updated financial information in real time, and greatly improves the efficiency of report production and update.

[0065] From Figure 3As can be seen, through the improved model, the present invention intuitively displays the annual financial data in the form of visual charts, aiming to improve the comprehensibility and operability of financial statements, facilitating financial analysis and management decision-making. The improved model combines various data presentation methods, including bar charts, line charts, pie charts, and tables, to comprehensively display core financial indicators such as revenue targets, actual revenues, and target completion rates for different months.

[0066] Figure 4 For the generated financial statements, the improved model can conduct multi-angle analysis of financial data, supporting management to carry out in-depth financial analysis and decision-making from a multi-dimensional perspective. Through the comprehensive presentation of various charts, managers can quickly identify the financial health status, discover potential problems, and formulate improvement measures.

[0067] Example 3

[0068] This example provides a financial statement automated processing device based on statistical analysis and image recognition, including:

[0069] Data acquisition unit: used to acquire the financial statement sample images to obtain the financial statement image data set;

[0070] Model construction unit: used to construct a financial statement target recognition network model based on the improved YOLOv8. The financial statement target recognition network model based on the improved YOLOv8 includes: Input input end, Backbone main network, Neck network, and Head output end. The images in the financial statement image data set are recognized by the financial statement target recognition network model based on the improved YOLOv8 to obtain the financial statement image text information;

[0071] Data analysis unit: used to perform pandas data analysis on the financial statement image text information through time series analysis methods to obtain the data analysis results, and generate relevant trend charts by using the matplotlib and seaborn libraries to achieve visualization;

[0072] Report generation unit: used to create a financial report in the Word document format based on the data analysis results and the generated relevant trend charts through the python-docx library.

[0073] Example 4

[0074] This embodiment provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the above-mentioned automated financial statement processing method based on statistical analysis and image recognition is executed.

[0075] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0076] Embodiment 5

[0077] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the above-mentioned automated financial statement processing method based on statistical analysis and image recognition is executed.

[0078] The storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0079] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for automated processing of financial statements based on statistical analysis and image recognition, characterized in that: The following steps are involved: S1. Obtain a sample image of a financial statement to obtain a financial statement image dataset; S2. Construct a financial statement target recognition network model based on improved YOLOv8, wherein the financial statement target recognition network model based on improved YOLOv8 includes: an input end, a backbone network, a neck network, and a head output end. The images in the financial statement image data set are recognized by the financial statement target recognition network model based on improved YOLOv8 to obtain financial statement image text information; The Backbone network includes a convolution Conv module, a C2 feature extraction layer C2f_1, a feature extraction layer C2f_Star module and an AIFI module; the convolution Conv module includes Conv_1, Conv_2, Conv_3, Conv_4, Conv_5 and Conv6, and the feature extraction layer C2f_Star module includes C2f_Star_1, C2f_Star_2 and C2f_Star_3; the C2f_Star_1 includes a convolution Conv1 module, 2 Star Blocks modules and a convolution Conv2 module; the Star Blocks module includes a depth-separable convolution DW-Conv1, 3 1×1 fully connected layers and a depth-separable convolution DW-Conv2, and the structures of C2f_Star_2 and C2f_Star_3 are the same as those of C2f_Star_1; The Neck network includes a feature-focused diffusion pyramid network FCF module, a convolutional layer Conv module, a C2 feature extraction layer, a Concat connection layer and an upsampled layer Upsample module; the feature-focused diffusion pyramid network FCF module includes FCF_1 and FCF_2, the convolutional layer Conv module includes Conv_6, Conv_7 and Conv_8, the C2 feature extraction layer includes C2f_2, C2f_3, C2f_4 and C2f_5, the Concat connection layer includes Concat_1, Concat_2, Concat_3 and Concat_4, and the upsampled layer Upsample module includes Upsample_1 and Upsample_2; S3. Perform pandas data analysis on the text information of financial statements through time series analysis method to obtain data analysis results. Use matplotlib and seaborn libraries to generate relevant trend charts from the data analysis results to achieve visualization; S4. Based on the results of data analysis and the generated relevant trend charts, the financial report in Word document format is created using the python-docx library.

2. The method for automated processing of financial statements based on statistical analysis and image recognition according to claim 1, characterized in that: The step S1 specifically includes: Set up an automated script, obtain the public financial statement sample data regularly based on the automated script, and remove the redundant information in the data to obtain the financial statement image data set .

3. The method for automated processing of financial statements based on statistical analysis and image recognition according to claim 2, characterized in that: The operation of the Input terminal in step S2 is specifically as follows: The Input input end includes a Mosaic data enhancement module, an adaptive anchor frame module and an adaptive image scaling module; Financial Statement Image Dataset The image in the input terminal is used to obtain the preprocessed financial statement image dataset , , Indicates A sample of preprocessed financial statement images.

4. The method for automated processing of financial statements based on statistical analysis and image recognition according to claim 3, characterized in that: The operations of the Backbone network in step S2 are as follows: Image samples from the preprocessed financial statement image dataset After the convolution layer Conv_1, the first feature map with 64 output channels is obtained. , the first feature map After the convolution layer Conv_2, the second feature map with 128 output channels is obtained. , the second feature map After the feature extraction layer C2f_1, the third feature map is obtained , the third feature map After the convolution layer Conv_3, the fourth feature map with 256 output channels is , the fourth characteristic map The fifth feature map with 256 output channels after the feature extraction layer C2f_Star_1 , the fifth characteristic map After the convolution layer Conv_4, the sixth feature map with 512 output channels is , the sixth characteristic map After the feature extraction layer C2f_Star_2, the seventh feature map with an output channel number of 512 is obtained , the seventh characteristic diagram After the convolution layer Conv_5, the output channel number is 1024, the eighth feature map , the eighth characteristic map After the feature extraction layer C2f_Star_3, the ninth feature map with an output channel number of 1024 is obtained , the ninth characteristic graph After downsampling in the convolutional layer Conv6, the number of channels becomes 512, and then it is input into the AIFI module to perform in-scale interaction on the high-level semantic features, and the tenth feature map with an output channel number of 512 is obtained. .

5. The method for automated processing of financial statements based on statistical analysis and image recognition according to claim 4, characterized in that: The specific operation of the Neck network in step S2 is: The seventh characteristic diagram , the tenth characteristic graph With the fifth characteristic diagram After the feature focused diffusion pyramid network FCF_1, the eleventh feature map with an output channel number of 128 is obtained , the eleventh feature map After the convolution layer Conv_6, the twelfth feature map with an output channel number of 256 is obtained , the twelfth characteristic graph With the tenth characteristic graph After the Concat connection layer Concat_1 is connected, the thirteenth feature map with an output channel number of 256 is obtained , the thirteenth characteristic diagram After the feature extraction layer C2f_2, the fourteenth feature map with an output channel number of 256 is obtained , the eleventh characteristic graph After the upsampling layer Upsample module Upsample_1, the upsampling is doubled, and the output is the fifteenth feature map with 256 output channels. , the fifteenth characteristic graph With the fifth characteristic diagram Through the Concat connection layer Concat_2 splicing, the sixteenth feature map with 256 output channels is obtained , the sixteenth characteristic diagram After the feature extraction layer C2f_3, the seventeenth feature map with 256 output channels is obtained. , the seventeenth characteristic diagram , Eleventh characteristic graph With the fourteenth characteristic figure After the feature focused diffusion pyramid network FCF_2, the eighteenth feature map with 256 output channels is obtained. , the eighteenth characteristic diagram After the Conv_7 module, the nineteenth feature map with an output channel number of 512 is obtained , the nineteenth characteristic diagram With the twelfth characteristic diagram 、The fourteenth characteristic graph Through the Concat_3 module, the twentieth feature map with an output channel number of 1024 is obtained , the twentieth feature map After the feature extraction layer C2f_4, the twenty-first feature map with an output channel number of 1024 is obtained , the twenty-first feature map After the upsampling layer Upsample module Upsample_2, the 22nd feature map is obtained , the fifteenth characteristic diagram 、22nd characteristic graph With the seventeenth characteristic figure After the Concat connection layer Concat_3, the twenty-third feature map with an output channel number of 1024 is obtained , the twenty-third feature map After the feature extraction layer C2f_5, the twenty-fourth feature map with an output channel number of 1024 is obtained .

6. The method for automated processing of financial statements based on statistical analysis and image recognition according to claim 5, characterized in that: The operation of the Head output terminal in step S2 is specifically as follows: The Head output end is composed of a Detect_DyHead module; the Detect_DyHead module includes a convolutional layer Conv module and a DyHeadBlock module; The eighteenth characteristic diagram 、The twenty-fourth characteristic diagram The detection frame is generated by the Detect_Dyhead detection module, the recognition of the financial statement image is completed, and the text information of the financial statement image is obtained; The financial statement image text information includes financial indicators, time stamps, market data and macroeconomic data.

7. The method for automated processing of financial statements based on statistical analysis and image recognition according to claim 1, characterized in that: Step S3 specifically includes: The data analysis results include the changing trends of financial indicators over time, cyclical fluctuations in the data, the correlation between different financial indicators and the correlation between macroeconomic indicators; the relevant trend charts include monthly income comparison chart, monthly target completion rate line chart, monthly completion rate percentage chart, monthly income comparison table, current ratio table, debt ratio table and profit rate table.

8. An automatic processing device for financial statements based on statistical analysis and image recognition, which executes the automatic processing method for financial statements based on statistical analysis and image recognition as claimed in claim 1, characterized in that: include: Data acquisition unit: used to acquire sample images of financial statements and obtain a financial statement image data set; Model building unit: used to build a financial statement target recognition network model based on improved YOLOv8, wherein the financial statement target recognition network model based on improved YOLOv8 includes: an input end, a backbone network, a neck network, and a head output end; images in the financial statement image data set are recognized by the financial statement target recognition network model based on improved YOLOv8 to obtain financial statement image text information; Data analysis unit: used to perform pandas data analysis on the text information of financial statements through time series analysis method, obtain data analysis results, and generate relevant trend charts from the data analysis results by using matplotlib and seaborn libraries to achieve visualization; Report generation unit: used to create financial reports in Word document format based on the results of data analysis and the generated related trend charts through the python-docx library.

9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the method for automated processing of financial statements based on statistical analysis and image recognition as described in any one of claims 1 to 7 is performed.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the automated processing method for financial statements based on statistical analysis and image recognition as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Road ponding recognition method, device and equipment and readable storage medium

    CN118570649A

  • Systems and methods for dental image analysis

    US20240046456A1