In my experince, I have found that the quality of the prompt is just as important as the quality of the rubric. When the prompt is clear, standardized, and explicitly aligned with the intended learning outcomes, the AI produces much more consistent and reliable evaluations. Rather than relying on generic instructions, you use structured prompts that require the AI to evaluate students work against predefined criteria and justify its feedback accordingly.
For auditing, I believe regular calibration against instructor evaluations is essential. Comparing AI outputs with your ratings on benchmark assignments, reviewing discrepancies, and refining the prompts when inconsistencies arise help maintain alignment with your standards. This combination of clear prompting, rubric-based evaluation, and periodic human review is really the most effective approach in my experience. I have two systems developed and works consistently.