score#

Functions

combine_gpu_arch_results(exec_result_dict, ...)

Combine results from different gpu archs by taking the worst correctness and average runtime.

combine_template_results(exec_result_list, ...)

Take the results for each parameter option and combine them into one result

compute_runtime_improvement(exec_result, ...)

Add improvement of kernel over baseline time to exec result, using ORM queries.

fastp(is_correct, baseline_speed, ...)

Rate of samples within a threshold p

geometric_mean_speed_ratio_correct_and_faster_only(...)

Geometric mean of the speed ratio for correct samples that have speedup > 1

geometric_mean_speed_ratio_correct_only(...)

Geometric mean of the speed ratio for correct samples

get_timing_stats(elapsed_times[, device, ...])

Get timing statistics from a list of elapsed times.

select_best_solution(eval_results)

Select the best kernel based on composite fitness score.

kernelfoundry.algorithm.utils.score.compute_runtime_improvement(exec_result, level: int, task_name: str, gpu_arch: str)[source]#

Add improvement of kernel over baseline time to exec result, using ORM queries.

Parameters:
  • exec_result – Execution result containing runtime but missing runtime_improvement

  • level – Level of the operation (e.g., 1)

  • task_name – Operation name (with .py!)

  • gpu_arch – gpu arch string

  • db – Database object with SessionRO() context manager

kernelfoundry.algorithm.utils.score.select_best_solution(eval_results: list[EvalResult]) → int[source]#

Select the best kernel based on composite fitness score.

Uses EvalResult.compute_performance_score to ensure consistent scoring across the system.

Parameters:

eval_results – List of EvalResult

Returns:

index of best solution

Return type:

int

kernelfoundry.algorithm.utils.score.combine_gpu_arch_results(exec_result_dict: dict[str, EvalResult], combined_eval_log: str) → EvalResult[source]#

Combine results from different gpu archs by taking the worst correctness and average runtime.

kernelfoundry.algorithm.utils.score.combine_template_results(exec_result_list: list, parameter_options: list, return_index: bool = False)[source]#

Take the results for each parameter option and combine them into one result

kernelfoundry.algorithm.utils.score.geometric_mean_speed_ratio_correct_only(is_correct: ndarray, baseline_speed: ndarray, actual_speed: ndarray, n: int) → float[source]#

Geometric mean of the speed ratio for correct samples

kernelfoundry.algorithm.utils.score.geometric_mean_speed_ratio_correct_and_faster_only(is_correct: ndarray, baseline_speed: ndarray, actual_speed: ndarray, n: int) → float[source]#

Geometric mean of the speed ratio for correct samples that have speedup > 1

kernelfoundry.algorithm.utils.score.fastp(is_correct: ndarray, baseline_speed: ndarray, actual_speed: ndarray, n: int, p: float) → float[source]#

Rate of samples within a threshold p

kernelfoundry.algorithm.utils.score.get_timing_stats(elapsed_times: list[float], device=None, device_name: str = 'unknown') → dict[source]#

Get timing statistics from a list of elapsed times.

Parameters:
  • elapsed_times – List of elapsed times in milliseconds

  • device – CUDA device, record device info

Returns:

Dict containing mean, std, min, max and num_trials all timing are in ms