Machine learning-based research and development cost proportion automatic calculation method and system
Through machine learning-based methods, the R&D expense related items are extracted from financial documents, the R&D expense ratio is calculated, and the R&D cost is regulated according to the development trend of the enterprise, the problem of unreasonable calculation of the R&D expense ratio in the existing technology is solved, and more scientific R&D cost management is achieved.
Patent Information
- Application Number
- CN202510034266.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-30
AI Technical Summary
It is difficult for existing technology to effectively calculate and regulate the proportion of R&D expenses of enterprises, resulting in unreasonable R&D budgets and affecting the long-term development strategy and project layout of enterprises.
Using a machine learning-based method, financial text features are obtained through financial-related documents, R&D expense related items are extracted, R&D expense ratio is calculated, and R&D costs are regulated using the Improved Ridge regression method according to the development trend of the enterprise.
The rapid, accurate and automated R&D expense ratio calculation has been achieved, and R&D costs can be reasonably regulated based on the current economic situation of the enterprise and improved R&D efficiency.
Smart Images

Figure CN120069980A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cost calculation, and particularly relates to a method and system for automatically calculating the R & D cost ratio based on machine learning. Background Art
[0002] Through the standardized research on R & D investment, problems such as resource waste and duplicate investment can be discovered and solved, and the quality and efficiency of enterprise R & D can be improved. The standardized research on R & D investment statistics also has important significance for enhancing international competitiveness. By aligning with international standards and establishing a scientific and standardized statistical system, it is possible to better participate in global scientific and technological cooperation and competition, and strengthen scientific and technological exchanges and cooperation with the international community.
[0003] At present, most enterprises have a large scale of R & D investment, involving various special projects of the comprehensive plan, self - arranged projects of each unit, and labor costs, etc. The coverage is wide, the number of coordinating departments is large, and the statistical difficulty is high. It is difficult to formulate a long - term development strategy that conforms to the actual situation, as well as a more reasonable R & D budget and project layout. Summary of the Invention
[0004] To solve the above problems existing in the prior art, the present invention provides a method and system for automatically calculating the R & D cost ratio based on machine learning.
[0005] The object of the present invention can be achieved by the following technical solutions:
[0006] A method for automatically calculating the R & D cost ratio based on machine learning, the implementation of the method for automatically calculating the R & D cost ratio includes the following steps:
[0007] Obtain financial text features through financial - related documents;
[0008] Extract R & D cost - related items according to the financial text features, and obtain the R & D cost ratio according to the R & D cost - related items;
[0009] Obtain the enterprise development trend according to the financial text features, and the enterprise development trend includes the positive development trend and the negative development trend of the enterprise;
[0010] When the negative development trend of the enterprise is output, the improved ridge regression method is used to control the enterprise R & D cost.
[0011] Preferably, the obtaining of the financial text features through financial - related documents includes:
[0012] Extract the text feature vector of the financial - related documents through the improved text extraction method;
[0013] Process the text feature vector through the auto - encoder model to obtain the financial text features.
[0014] Preferably, the extraction of the text feature vector of the financial-related document by the improved text extraction method includes:
[0015] Removing irrelevant characters from the financial-related document;
[0016] Using the jieba library in Python and the Harbin Institute of Technology stop word list to remove invalid words from the financial-related document;
[0017] Processing the financial-related document through the BERT model and obtaining the text feature vector.
[0018] Preferably, the processing of the text feature vector by the autoencoder model to obtain the financial text feature includes:
[0019] Mapping the text feature vector to the latent representation space, and the mathematical description of the latent representation space is z = g(x) = σ(Wx + b), where z is the latent representation space, g(x) is the autoencoder function, σ is the activation function, W is the autoencoder weight matrix, x is the text feature vector, and b is the autoencoder bias vector;
[0020] Mapping the latent representation space to the output layer and performing decoding and reconstruction through the decoder to obtain the financial text feature, and the mathematical description of the financial text feature is Z = f(z) = σ(W T z + B), where Z is the financial text feature, f(z) is the decoder function, z is the latent representation space, W T is the transpose matrix of the autoencoder weight matrix, and B is the decoder bias vector.
[0021] Preferably, extracting the R & D expense-related items according to the financial text feature and obtaining the R & D expense ratio according to the R & D expense-related items includes:
[0022] Retrieving the R & D expense-related items in the financial text feature;
[0023] Automatically calculating the R & D expense ratio according to the R & D expense-related items, and the R & D expense ratio is the ratio of each R & D expense-related item to the total R & D expense.
[0024] Preferably, obtaining the enterprise development trend according to the financial text feature includes:
[0025] Presetting enterprise development trend words, and the enterprise development trend words include enterprise positive development words and enterprise negative development words;
[0026] Calculating the occurrence frequency of the enterprise development trend words in the financial text feature, and the occurrence frequency includes the occurrence frequency of enterprise positive development words and the occurrence frequency of enterprise negative development words;
[0027] Calculate the inverse document frequency, which includes the positive inverse document frequency and the negative inverse document frequency;
[0028] Calculate the high word frequency according to the occurrence frequency and the inverse document frequency. The high word frequency includes the positive high word frequency and the negative high word frequency. The calculation formula for the positive high word frequency is γ 1 =A 1 ×IDF 1 , where γ 1 is the positive high word frequency, A 1 is the occurrence frequency of the enterprise's positive development words, IDF 1 is the positive inverse document frequency. The calculation formula for the negative high word frequency is γ 2 =A 2 ×IDF 2 , where γ 2 is the negative high word frequency, A 2 is the occurrence frequency of the enterprise's negative development words, IDF 2 is the negative inverse document frequency;
[0029] Obtain the enterprise development trend according to the high word frequency. When the positive high word frequency is greater than or equal to the negative high word frequency, output the positive development trend of the enterprise. When the positive high word frequency is less than the negative high word frequency, output the negative development trend of the enterprise.
[0030] Preferably, the regulation of the enterprise's R & D cost by using the improved ridge regression method includes:
[0031] Obtain the enterprise R & D labels, which include R & D project type labels, R & D project level labels, R & D project difficulty labels, R & D undertaking department labels, R & D project number labels, and R & D project cycle labels. The R & D project type labels include new R & D project labels, imitation R & D project labels, and improved R & D project labels. The R & D project level labels include general R & D project labels and key R & D project labels. The R & D project difficulty labels include easy R & D project labels, medium R & D project labels, and difficult R & D project labels. The R & D undertaking department labels include R & D project labels of the design department, R & D project labels of the R & D department, and R & D project labels of the production technology department;
[0032] Obtain the historical R & D cost data, and obtain the strongly related enterprise R & D labels through the historical R & D cost data. The strongly related enterprise R & D labels are the top five enterprise R & D labels ranked by the correlation with the historical R & D cost data. The strongly related enterprise R & D labels include the first strongly related enterprise R & D label, the second strongly related enterprise R & D label, the third strongly related enterprise R & D label, the fourth strongly related enterprise R & D label, and the fifth strongly related enterprise R & D label;
[0033] Obtain the estimated coefficients through ridge regression. The estimated coefficients include the zero - th estimated coefficient, the first estimated coefficient, the second estimated coefficient, the third estimated coefficient, the fourth estimated coefficient, and the fifth estimated coefficient;
[0034] Obtain the estimated coefficient matrix;
[0035] Calculate the estimated coefficients. The calculation formula is ε i =(MM T +K i I) -1 M T Y, where i = 0, 1, 2,..., 5. Here, ε i is the i - th estimated coefficient, M is the estimated coefficient matrix, K i is the ridge parameter, with a value in the closed interval [0.1, 0.3], I is the identity matrix, and Y is the enterprise R & D cost;
[0036] Construct an enterprise R & D cost function based on the estimated coefficients;
[0037] Regulate the enterprise R & D cost according to the enterprise R & D cost function.
[0038] A machine - learning - based automatic calculation system for R & D expense ratio, which is used to execute the above - mentioned automatic calculation method for R & D expense ratio. It is characterized by including an R & D expense calculation module, an enterprise development trend judgment module, and an enterprise R & D cost regulation module;
[0039] The R & D expense calculation module is used to obtain financial text features through financial - related documents, extract R & D expense - related items according to the financial text features, and obtain the R & D expense ratio according to the R & D expense - related items;
[0040] The enterprise development trend judgment module is used to obtain the enterprise development trend according to the financial text features. The enterprise development trend includes an enterprise positive development trend and an enterprise negative development trend;
[0041] The enterprise R & D cost regulation module is used to regulate the enterprise R & D cost by using the improved ridge regression method when the enterprise negative development trend is output.
[0042] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above - mentioned automatic calculation method for R & D expense ratio.
[0043] A storage medium containing computer - executable instructions, and the computer - executable instructions are used to execute the above - mentioned automatic calculation method for R & D expense ratio when executed by a computer processor.
[0044] The beneficial effects of the present invention are as follows:
[0045] (1) Obtain financial text features from financial-related documents, extract items related to research and development expenses based on the financial text features, and obtain the research and development expense ratio based on the items related to research and development expenses, so as to achieve rapid, accurate, and automated calculation of the research and development expense ratio;
[0046] (2) Obtain the enterprise development trend through financial text features, and adopt the improved ridge regression method to regulate the enterprise's research and development costs. The regulation process conforms to the enterprise's economic status, and can more reasonably conduct scientific regulation on different enterprises;
[0047] (3) Process the text feature vector through an autoencoder model to obtain financial text features, use a neural network to compress the input data into a low-dimensional space, and remap the compressed data back to the input space through a decoder for reconstruction, so as to achieve the purpose of reducing the data dimension and avoid the problem of dimensional disaster caused by too high text information dimension;
[0048] (4) Calculate the high word frequency according to the occurrence frequency and inverse document frequency, obtain the enterprise development trend according to the high word frequency, filter out common words, and retain important words that can provide more information about the company's operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] For the convenience of those skilled in the art to understand, the present invention will be further described below with reference to the accompanying drawings.
[0050] Figure 1 It is a flowchart of the steps of an automatic calculation method for the research and development expense ratio based on machine learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will describe in detail the specific embodiments, structures, features, and their effects of the present invention with reference to the accompanying drawings and preferred embodiments.
[0052] The working principle and usage process of the present invention:
[0053] Please refer to Figure 1 , an automatic calculation method for the research and development expense ratio based on machine learning, including:
[0054] S1: Obtain financial text features from financial-related documents, where the financial-related documents include but are not limited to company annual reports, company research reports, and financial reports;
[0055] S2: Extract items related to research and development expenses according to the financial text features, and obtain the research and development expense ratio according to the items related to research and development expenses;
[0056] S3: Obtain the enterprise development trend according to the financial text features, where the enterprise development trend includes the positive development trend of the enterprise and the negative development trend of the enterprise;
[0057] S4: When outputting the negative development trend of the enterprise, use the improved ridge regression method to regulate the R & D cost of the enterprise.
[0058] In this embodiment, the financial text features are obtained through financial-related documents, and the specific implementation can be carried out through the following steps:
[0059] S101: Extract the text feature vector of the financial-related document through the improved text extraction method;
[0060] S102: Process the text feature vector through an autoencoder model to obtain the financial text features. Use a neural network to compress the input data into a low-dimensional space, and then remap the compressed data back to the input space through a decoder for reconstruction, so as to achieve the purpose of reducing the data dimension and avoid the problem of dimensional disaster caused by too high text information dimension.
[0061] In this embodiment, the text feature vector of the financial-related document is extracted through the improved text extraction method, and the specific implementation can be carried out through the following steps:
[0062] S101-1: Remove the irrelevant characters in the financial-related document, where the irrelevant characters include but are not limited to punctuation marks, letters, and spaces;
[0063] S101-2: Use the jieba library in Python and the Harbin Institute of Technology stop word list to remove the invalid words in the financial-related document, where the invalid words include but are not limited to "ah", "ba", "ya";
[0064] S101-3: Process the financial-related document through the BERT model and obtain the text feature vector. The processing process of the BERT model is as follows: Divide the input financial-related document into several tokens and map each token to a word vector, use multiple bidirectional Transformer encoders to encode the word vector to obtain a feature vector, and finally converge the feature vector into a text feature vector.
[0065] In this embodiment, the financial text features are obtained by processing the text feature vector through an autoencoder model, and the specific implementation can be carried out through the following steps:
[0066] S102-1: Map the text feature vector to the latent representation space. The mathematical description of the latent representation space is z = g(x) = σ(Wx + b), where z is the latent representation space, g(x) is the autoencoder function, σ is the activation function, W is the autoencoder weight matrix, x is the text feature vector, and b is the autoencoder bias vector;
[0067] S102-2: Map the potential representation space to the output layer and perform decoding and reconstruction through a decoder to obtain the financial text feature. The mathematical description of the financial text feature is Z = f(z) = σ(W T z + B), where Z is the financial text feature, f(z) is the decoder function, z is the potential representation space, and W T is the transpose matrix of the autoencoder weight matrix, and B is the decoder bias vector.
[0068] In this embodiment, relevant items related to R & D expenses are extracted from the financial text feature, and the R & D expense ratio is obtained based on the relevant items related to R & D expenses. Specifically, it can be implemented through the following steps:
[0069] S201: Retrieve the relevant items related to R & D expenses in the financial text feature. The relevant items related to R & D expenses include but are not limited to salaries and benefits of R & D personnel, depreciation expenses of R & D equipment, repair expenses of R & D equipment, maintenance expenses of R & D equipment, research materials and manufacturing expenses, buildings and rents for research and development, research and test expenses, intellectual property expenses such as patents, copyrights and trademarks, and other expenses related to research and development;
[0070] S202: Automatically calculate the R & D expense ratio based on the relevant items related to R & D expenses. The R & D expense ratio is the ratio of each relevant item related to R & D expenses to the total R & D expenses.
[0071] In this embodiment, the enterprise development trend is obtained based on the financial text feature. The enterprise development trend includes the positive development trend of the enterprise and the negative development trend of the enterprise. Specifically, it can be implemented through the following steps:
[0072] S301: Preset enterprise development trend words, which include positive enterprise development words and negative enterprise development words. The positive enterprise development words include but are not limited to applicable, develop, improve, profit, develop, advantage, enhance, and the negative enterprise development words include but are not limited to reduce, risk, adverse, loss, decline, downward;
[0073] S302: Calculate the occurrence frequency of the enterprise development trend words in the financial text feature. The occurrence frequency includes the occurrence frequency of positive enterprise development words and the occurrence frequency of negative enterprise development words. The calculation formula for the occurrence frequency of positive enterprise development words is where A 1 is the occurrence frequency of positive enterprise development words, n 1 is the number of occurrences of positive enterprise development words, and n is the total number of occurrences of all words in the financial text feature. The calculation formula for the occurrence frequency of negative enterprise development words is where A 2It is the frequency of occurrence of negative development words for the enterprise, n 2 It is the number of occurrences of negative development words for the enterprise;
[0074] S303: Calculate the inverse document frequency, where the inverse document frequency includes a positive inverse document frequency and a negative inverse document frequency. The calculation formula for the positive inverse document frequency is where IDF 1 is the positive inverse document frequency, N is the number of documents in the financial-related documents, and u 1 is the number of documents in the financial-related documents that contain positive development words for the enterprise. The calculation formula for the negative inverse document frequency is where IDF 2 is the negative inverse document frequency, and u 2 is the number of documents in the financial-related documents that contain negative development words for the enterprise;
[0075] S304: Calculate the high word frequency according to the frequency of occurrence and the inverse document frequency. The high word frequency includes a positive high word frequency and a negative high word frequency. The calculation formula for the positive high word frequency is γ 1 = A 1 × IDF 1 , where γ 1 is the positive high word frequency, A 1 is the frequency of occurrence of positive development words for the enterprise, and IDF 1 is the positive inverse document frequency. The calculation formula for the negative high word frequency is γ 2 = A 2 × IDF 2 , where γ 2 is the negative high word frequency, A 2 is the frequency of occurrence of negative development words for the enterprise, and IDF 2 is the negative inverse document frequency;
[0076] S305: Obtain the development trend of the enterprise according to the high word frequency, filter out common words, and retain important words that can provide more information about the company's operation. When the positive high word frequency is greater than or equal to the negative high word frequency, output the positive development trend of the enterprise. When the positive high word frequency is less than the negative high word frequency, output the negative development trend of the enterprise.
[0077] In this embodiment, when outputting the negative development trend of the enterprise, the improved ridge regression method is used to regulate the R & D cost of the enterprise, which can be specifically implemented through the following steps:
[0078] S401: Obtain enterprise R & D labels, where the enterprise R & D labels include R & D project type labels, R & D project level labels, R & D project difficulty labels, R & D undertaking department labels, R & D project number labels, and R & D project cycle labels. The R & D project type labels include new R & D project labels, imitation R & D project labels, and improvement R & D project labels. The R & D project level labels include general R & D project labels and key R & D project labels. The R & D project difficulty labels include easy R & D project labels, medium R & D project labels, and difficult R & D project labels. The R & D undertaking department labels include R & D project labels of the design department, R & D project labels of the R & D department, and R & D project labels of the production technology department;
[0079] S402: Obtain historical R & D expense data, and extract the strongly correlated enterprise R & D labels with the highest correlation with the historical R & D expense data. The strongly correlated enterprise R & D labels include the first strongly correlated enterprise R & D label, the second strongly correlated enterprise R & D label, the third strongly correlated enterprise R & D label, the fourth strongly correlated enterprise R & D label, and the fifth strongly correlated enterprise R & D label;
[0080] S403: Obtain the estimated coefficients through ridge regression. The estimated coefficients include the zero - th estimated coefficient, the first estimated coefficient, the second estimated coefficient, the third estimated coefficient, the fourth estimated coefficient, and the fifth estimated coefficient;
[0081] S404: Obtain the estimated coefficient matrix, where the estimated coefficient matrix is the matrix composed of the estimated coefficients;
[0082] S405: Calculate the estimated coefficients. The calculation formula is ε i =(MM T +K i I) -1 M T Y, i = 0, 1, 2,..., 5, where ε i is the i - th estimated coefficient, M is the estimated coefficient matrix, K i is the ridge parameter, with a value in the closed interval [0.1, 0.3], I is the identity matrix, and Y is the enterprise R & D cost;
[0083] S406: Construct an enterprise R & D cost function according to the estimated coefficients. The expression of the enterprise R & D cost function is Y = ε 0 +ε 1 L 1 +ε 2 L 2 +ε 3 L 3 +ε 4 L 4 +ε 5 L 5 , where Y is the enterprise R & D cost, εi is the i-th estimated coefficient, where i = 0, 1, 2, …, 5, L j is the R & D label of the j-th strongly correlated enterprise, where j = 1, 2, …, 5;
[0084] S407: Regulate the R & D cost of the enterprise according to the enterprise R & D cost function, and preferentially regulate the R & D label of the strongly correlated enterprise with a larger estimated coefficient.
[0085] An automatic calculation system for R & D expense ratio based on machine learning, comprising an R & D expense calculation module, an enterprise development trend judgment module, and an enterprise R & D cost regulation module;
[0086] The R & D expense calculation module is used to obtain financial text features through financial related documents, extract R & D expense related items according to the financial text features, and obtain the R & D expense ratio according to the R & D expense related items;
[0087] The enterprise development trend judgment module is used to obtain the enterprise development trend according to the financial text features, and the enterprise development trend includes an enterprise positive development trend and an enterprise negative development trend;
[0088] The enterprise R & D cost regulation module is used to regulate the enterprise R & D cost by using the improved ridge regression method when the enterprise negative development trend is output.
[0089] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0090] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0091] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing. The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0092] As described above, the above are only the preferred embodiments of the present invention and do not impose any formal limitations on the present invention. Although the present invention has been disclosed as above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or refinements to the equivalent embodiments with equivalent changes within the scope of the technical solution of the present invention. However, any simple modification, equivalent change, and refinement made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A method for automatically calculating the ratio of R&D expenses based on machine learning, characterized in that: The implementation of the automatic calculation method of the R&D expense ratio includes the following steps: Obtain financial text features through financial related documents; Extracting R&D expense related items according to the financial text features, and obtaining the R&D expense ratio according to the R&D expense related items; Acquire the enterprise development trend according to the financial text feature, wherein the enterprise development trend includes a positive enterprise development trend and a negative enterprise development trend; When the output shows a negative development trend for the enterprise, the improved ridge regression method is used to regulate the enterprise's R&D costs.
2. The method for automatically calculating the ratio of R&D expenses according to claim 1, characterized in that: The obtaining of financial text features through financial-related documents includes: Extracting text feature vectors of the financial-related documents by using an improved text extraction method; The financial text feature is obtained by processing the text feature vector through an autoencoder model.
3. The method for automatically calculating the ratio of R&D expenses according to claim 2, characterized in that: The step of extracting the text feature vector of the financial-related document by using the improved text extraction method comprises: Removing irrelevant characters from said financial-related documents; Use the jieba library in Python and the Harbin Institute of Technology stop word list to remove invalid words in the financial-related documents; The financial-related document is processed through a BERT model to obtain the text feature vector.
4. The method for automatically calculating the ratio of R&D expenses according to claim 2, characterized in that: Processing the text feature vector by the autoencoder model to obtain the financial text feature includes: Mapping the text feature vector to a latent representation space, the mathematical description of the latent representation space is z=g(x)=σ(Wx+b), where z is the latent representation space, g(x) is the autoencoder function, σ is the activation function, W is the autoencoder weight matrix, x is the text feature vector, and b is the autoencoder bias vector; The latent representation space is mapped to the output layer, and the financial text feature is obtained by decoding and reconstructing through a decoder. The mathematical description of the financial text feature is Z = f(Z) = σ(WTz+B), where Z is the financial text feature, f(z) is the decoder function, z is the latent representation space, and W T is the transposed matrix of the autoencoder weight matrix, and B is the decoder bias vector.
5. The method for automatically calculating the ratio of R&D expenses according to claim 1, characterized in that: The extracting R&D expense related items according to the financial text features and obtaining the R&D expense ratio according to the R&D expense related items include: Retrieving the R&D expense related items in the financial text features; The R&D expense ratio is automatically calculated based on the R&D expense related items, and the R&D expense ratio is the ratio of each R&D expense related item to the total R&D expense.
6. The method for automatically calculating the ratio of R&D expenses according to claim 1, characterized in that: The obtaining of the enterprise development trend according to the financial text features includes: Preset enterprise development trend words, the enterprise development trend words include enterprise positive development words and enterprise negative development words; Calculating the occurrence frequency of the enterprise development trend words in the financial text features, wherein the occurrence frequency includes the occurrence frequency of enterprise positive development words and the occurrence frequency of enterprise negative development words; Calculating reverse document frequencies, wherein the reverse document frequencies include positive reverse document frequencies and negative reverse document frequencies; The high word frequency is calculated according to the occurrence frequency and the reverse document frequency, wherein the high word frequency includes positive high word frequency and negative high word frequency, and the calculation formula of the positive high word frequency is γ1=A1×IDF1, wherein γ1 is the positive high word frequency, A1 is the occurrence frequency of the enterprise's positive development words, and IDF1 is the positive reverse document frequency. The calculation formula of the negative high word frequency is γ2=A2×IDF2, wherein γ2 is the negative high word frequency, A2 is the occurrence frequency of the enterprise's negative development words, and IDF2 is the negative reverse document frequency; The enterprise development trend is obtained according to the high word frequency. When the positive high word frequency is greater than or equal to the negative high word frequency, the positive enterprise development trend is output. When the positive high word frequency is less than the negative high word frequency, the negative enterprise development trend is output.
7. The method for automatically calculating the ratio of R&D expenses according to claim 1, characterized in that: The improved ridge regression method is used to regulate the R&D costs of enterprises, including: Obtain enterprise R&D labels, wherein the enterprise R&D labels include R&D project type labels, R&D project level labels, R&D project difficulty labels, R&D undertaking department labels, R&D project headcount labels, and R&D project cycle labels. The R&D project type labels include new R&D project labels, imitation R&D project labels, and improved R&D project labels. The R&D project level labels include general R&D project labels and key R&D project labels. The R&D project difficulty labels include easy R&D project labels, medium R&D project labels, and difficult R&D project labels. The R&D undertaking department labels include design department R&D project labels, R&D department R&D project labels, and production technology department R&D project labels. Obtain historical R&D expense data, and obtain strongly relevant enterprise R&D labels through the historical R&D expense data, wherein the strongly relevant enterprise R&D labels are the enterprise R&D labels ranked in the top five in terms of relevance to the historical R&D expense data, and the strongly relevant enterprise R&D labels include the first strongly relevant enterprise R&D label, the second strongly relevant enterprise R&D label, the third strongly relevant enterprise R&D label, the fourth strongly relevant enterprise R&D label, and the fifth strongly relevant enterprise R&D label; Obtain estimated coefficients by ridge regression, wherein the estimated coefficients include a zeroth estimated coefficient, a first estimated coefficient, a second estimated coefficient, a third estimated coefficient, a fourth estimated coefficient, and a fifth estimated coefficient; Get the estimated coefficient matrix; Calculate the estimated coefficient, the calculation formula is ε i =(MM T +K i I) -1 M T Y,i=0,1,2,...,,5, where,ε i is the i-th estimated coefficient, M is the estimated coefficient matrix, K i is the ridge parameter, with a value of the closed interval [0.1, 0.3], I is the unit matrix, and Y is the R&D cost of the enterprise; Constructing the enterprise R&D cost function according to the estimated coefficients; The enterprise R&D cost is regulated according to the enterprise R&D cost function.
8. A system for automatically calculating the proportion of R&D expenses, characterized in that: The system is applied to the automatic calculation method of the ratio of R&D expenses as described in any one of claims 1 to 7, including an R&D expense calculation module, an enterprise development trend judgment module, and an enterprise R&D cost control module; The R&D expense calculation module is used to obtain financial text features through financial related documents, extract R&D expense related items according to the financial text features, and obtain the R&D expense ratio according to the R&D expense related items; The enterprise development trend judgment module is used to obtain the enterprise development trend according to the financial text features, and the enterprise development trend includes a positive enterprise development trend and a negative enterprise development trend; The enterprise R&D cost control module is used to control the enterprise R&D cost by adopting the improved ridge regression method when the negative development trend of the enterprise is output.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for automatically calculating the ratio of R&D expenses as described in any one of claims 1-7 is implemented.
10. A storage medium containing computer executable instructions, characterized in that: The computer executable instructions, when executed by a computer processor, are used to execute the method for automatically calculating the ratio of R&D expenses as described in any one of claims 1 to 7.