The invention provides an improved Q-learning
channel access method based on a composite reward function, and belongs to the technical field of communication. Comprising the following steps: an initialization
algorithm: deploying sensor nodes and gateways, clustering and selecting cluster heads from nodes in clusters; locally initializing nodes, maintaining a Q table and a neighbor table by each node, and defining a
state space, an action space, an initial parameter and a competition intensity threshold value; each node executes actions and observes in the current state, when a conflict does not pass in the data sending process, a composite reward participating in Q value updating is calculated, layered Q value updating based on the competitive strength level is firstly carried out after updating to the next state, a neighbor table is updated by the node in the Texchange period, and the neighbor table is updated by the node in the Texchange period; and in the Tcluster period, the weighted average Q value in the cluster is calculated by the cluster head and is broadcasted, and after the cluster member updates the local Q table, the action is executed and the observation is carried out in the next state until the conflict is solved and the node information is successfully sent. According to the method, the conflict probability in a high-competition environment can be effectively reduced, and the channel access efficiency is improved.